Aetherfy limits
What this Aetherfy limits page covers
This page is the Aetherfy platform limit index: quotas that belong to your account and plan. Per-request caps on the vector API — batch sizes, payload sizes, result counts — are documented at Vector database limits and are not repeated here.
Aetherfy plan limits by plan
| Limit | Free | Starter | Performance | Enterprise |
|---|---|---|---|---|
| Agents | 1 | 3 | 10 | unlimited |
| Max memory per agent | 256 MB | 1 GB | 2 GB | 8 GB |
| Regions | 1 | 1 | 3 | unlimited |
| Always-on allowed | no | yes | yes | yes |
| Collections | 3 | 30 | 200 | unlimited |
| API keys | 2 | 5 | 10 | 50 |
| Workspaces | 1 | unlimited | unlimited | unlimited |
| Storage | 512 MB | 10 GB | 200 GB | unlimited |
| Vector API requests/min | 1 000 | 5 000 | 20 000 | 50 000 |
| Control-plane API requests/min | 100 | 500 | 2 000 | unlimited |
| Custom Dockerfile builds | no | yes | yes | yes |
| Max idle timeout | 5 min | 15 min | 30 min | unlimited |
| Runs in flight | 1 | 10 | 50 | negotiated |
Where an Aetherfy limit reads NULL or “unlimited”, no limit is enforced for
that plan.
Note the region row: the Free and Starter plans of Aetherfy are single-region, and multi-region placement begins at Performance. See Regions and replication.
Note the collection row, because of what it does NOT count: memory threads do not consume collections. Every thread in a workspace is stored in one shared collection keyed by the thread id, so the number of conversations you can hold is not limited by this row — only a namespace costs a collection. See The memory API.
Note the runs-in-flight row, because it is the one limit on this page that is
about ACTIVITY rather than about what you own. It counts task runs that are
executing at this moment, across every task agent on the account and however
each run was triggered — a run you started by hand, a run a schedule fired, and
a run one agent spawned in another all count the same. It is the only limit on
how many runs of one agent execute at once: a run that finds its agent’s machine
busy gets another machine. Start one past the limit by hand or by spawning and
Aetherfy refuses it with 429 AGENT_RUN_CONCURRENCY_LIMIT_EXCEEDED, whose
detail.limit reads max_in_flight_runs; a scheduled occurrence past it is
skipped and recorded with that same code. Nothing is queued, and no run row is
created. A service run adds no machine and is not counted. Runs that have finished do not count, so the limit is about how wide
you fan out at once and not about how much you run in a day. Enterprise is
negotiated per contract rather than unlimited by default.
Note the max-idle-timeout row too, because Aetherfy enforces it in one layer
only: the per-plan maximum applies when an agent is created or updated through
the Aetherfy API or the dashboard. The aetherfy.yaml parser and the deploy
path place no upper bound on idle_timeout_minutes — see
the aetherfy.yaml reference.
Aetherfy has two separate rate limits
This corrects previously published documentation, so it is worth stating plainly: Aetherfy enforces two independent per-minute rate limits, not one.
| Plane | Host | Limit |
|---|---|---|
| Vector data plane | vectors.aetherfy.com | Vector API requests/min from the table above |
| Control plane | agents.aetherfy.com | Control-plane API requests/min from the table above |
The two budgets are separate. Saturating the control plane does not consume vector
API allowance, and vice versa. Exceeding either returns HTTP 429 carrying the code
RATE_LIMIT_EXCEEDED — though the two planes wrap that code in different envelopes,
covered below.
How the Aetherfy rate-limit window works
The two Aetherfy planes do not share a limiter implementation, and the difference is visible to a client. Write your backoff against the plane you are calling.
Vector data plane (vectors.aetherfy.com) | Control plane (agents.aetherfy.com) | |
|---|---|---|
| Window | Fixed 60-second bucket | Sliding 60-second window |
Retry-After on a 429 | Not sent | Sent |
X-RateLimit-* headers | Not sent | Sent — Limit, Remaining, Reset |
| 429 envelope | error.code | detail.code |
On the vector data plane the counter resets at the bucket boundary rather than
gradually releasing capacity, so a burst that straddles a boundary can succeed where
the same burst inside one bucket would not. There is no Retry-After header on a
vector-plane 429 — back off on your own schedule, and exponential backoff with jitter
is the sane default.
On the control plane the window slides: capacity is returned continuously as individual requests age past 60 seconds, so there is no boundary to straddle. A control-plane 429 tells you exactly when to come back, and you should read it rather than guess:
HTTP/1.1 429 Too Many Requests
Retry-After: 23
X-RateLimit-Limit: 500
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1755600000{
"detail": {
"code": "RATE_LIMIT_EXCEEDED",
"message": "Rate limit exceeded. Limit: 500 requests/minute. Try again in 23 seconds.",
"limit": 500,
"retry_after_seconds": 23
}
}Aetherfy also returns X-RateLimit-Limit, X-RateLimit-Remaining and
X-RateLimit-Reset on successful control-plane responses, so a client can pace
itself without ever provoking a 429. X-RateLimit-Reset is a Unix timestamp in
seconds.
Enterprise accounts have no control-plane request limit, and Aetherfy sends no
X-RateLimit-* headers at all on that plan — an unlimited budget has no remaining
count to report. Treat the headers as optional rather than guaranteed.
This corrects earlier published guidance which stated the fixed-bucket window and the
absence of Retry-After as platform-wide facts. Both were only ever true of the
vector data plane.
How Aetherfy counts storage
Aetherfy storage quota counts every replica. A collection’s bytes are counted once per region it is placed in, so 100 GB held in one region and 33 GB replicated across three regions consume the same storage budget.
Exceeding the storage limit blocks writes. Reads are never blocked — an Aetherfy account over its storage cap can still serve search and retrieval traffic while you free space or move to a larger plan.
Aetherfy agent memory values
Agent memory in Aetherfy must be one of these exact values, in megabytes, and must also be within your plan’s maximum from the table above:
| Allowed memory values (MB) |
|---|
256 |
512 |
1024 |
2048 |
8192 |
An arbitrary value between these is not accepted.
Cores follow memory. Aetherfy gives an agent one shared vCPU per 2048 MB of
memory, so the 8192 step runs on four shared vCPUs and every smaller step on
one. There is no separate core setting to choose.
Fixed Aetherfy limits that no plan changes
| Limit | Value |
|---|---|
| Task run duration | Terminated at 60 minutes and recorded as failed |
| Log retention | 7 days |
Neither is configurable on any Aetherfy plan, including Enterprise. A workload that legitimately needs more than 60 minutes should be split into several scheduled task runs that checkpoint their progress.
How Aetherfy counts agents against your limit
| Agent state | Consumes an agent slot |
|---|---|
| Created, never deployed | No |
| Deployed | Yes |
| Paused | Yes |
| Archived | No |
The limit counts deployed agents, so creating one is never refused for being over it — a created agent is a draft until you deploy it, and drafts are not rationed on any plan. A paused Aetherfy agent still holds its slot, so pausing agents is not a way to fit more of them under your plan’s agent limit. Archiving is.
What Aetherfy returns when a plan limit is exceeded
Requests that would take you past a plan limit are rejected with HTTP 403 and
the code PLAN_LIMIT_EXCEEDED. The message names the limit that fired and what
would resolve it. It covers the agent count, memory per agent, region count,
always-on, the idle timeout, and custom Dockerfile builds — the whole plan-cap
family shares this one code, so branch on the message rather than expecting a
distinct code per limit.
This is separate from the usage-limit codes: PLAN_LIMIT_EXCEEDED means a
structural cap on your plan, while SOFT_CAP_EXCEEDED means you are within your
plan but at your spend limit. See Billing and spend caps.
Aetherfy limits FAQ
Which rate limit applies to my request?
It depends on which Aetherfy host you called. Requests to the vector data plane
at vectors.aetherfy.com — searching, upserting, retrieving and deleting points
— draw on the vector API allowance, which is 1 000 requests per minute on Free,
5 000 on Starter, 20 000 on Performance and 50 000 on Enterprise. Requests to the
control plane at agents.aetherfy.com — deploying, spawning, reading logs,
managing workspaces — draw on a completely separate control-plane allowance of
100, 500, 2 000 and unlimited respectively. These are two independent budgets:
exhausting one leaves the other untouched.
Does storage count once, or once per region?
Once per region. Aetherfy counts every replica against your storage quota, so a collection replicated into three regions consumes three times its own size. This is why 100 GB in a single region and 33 GB across three regions cost you the same budget. It also means that adding a region to a large existing collection is a storage decision as much as a latency one — check your headroom before you widen a collection’s placement.
What happens if I exceed my storage cap?
Aetherfy blocks writes and leaves reads alone. Upserts and other write operations are refused while you are over the cap, but search, retrieval and every other read path continue to serve normally. That asymmetry is deliberate: an application over its storage limit degrades to read-only rather than going dark. To recover, delete vectors you no longer need, release a replica in a region you no longer serve, or move to a plan with a larger storage allowance.
Do paused or archived agents count against my agent limit?
Paused agents do; archived agents do not. Pausing an Aetherfy agent stops it from doing work but keeps its slot reserved, so you cannot pause your way under the agent limit. Archiving releases the slot. If you are at your agent limit and want to deploy something new without upgrading, archive an agent you are not using rather than pausing it.
Is there a Retry-After header on a 429?
It depends on the Aetherfy plane. The control plane at agents.aetherfy.com
does send Retry-After, alongside X-RateLimit-Limit, X-RateLimit-Remaining and
X-RateLimit-Reset, and its 429 body repeats the wait as retry_after_seconds —
read it rather than guessing. The vector data plane at vectors.aetherfy.com
sends none of those, so there you must implement your own backoff, with exponential
plus jitter as the standard choice. The window differs too: the vector plane is a
fixed 60-second bucket, so waiting out the remainder of the current minute recovers
the full allowance, while the control plane slides and returns capacity continuously.
Earlier documentation stated the vector-plane behaviour as though it applied to both;
it does not.
How long are logs kept?
Aetherfy retains logs for 7 days, on every plan. This is not configurable, and no plan extends it. If you need longer retention — for audit, compliance or trend analysis — ship the logs to your own destination from within your agent while they are still inside the 7-day window.