Skip to Content
Agent computeThe task contract
Raw

The task contract

How Aetherfy runs your code

A task on Aetherfy runs as a plain script. There is no framework to import, no handler function to export, and no server to start. Aetherfy resumes or starts your machine, executes your entrypoint file, and reads the exit code when the process ends.

Your entrypoint is executed top to bottom, exactly as if you had run it yourself. A Python if __name__ == "__main__": block fires normally, and a Node file’s top-level statements run normally.

Each run is a fresh process. Nothing your script leaves in memory is visible to the next run, so treat every run as starting from nothing — the machine may have been asleep for a month or for a second, and your code cannot tell the difference.

This applies to type: job agents. Long-lived service agents on Aetherfy serve HTTP instead; the one thing a scheduled or manual run asks of a service is covered in its own section below. It applies to runtime: dockerfile too: a task in a custom container runs the command your image declares, once per run, and everything on this page holds for it — see Custom Dockerfile runtime.

Exit codes and how Aetherfy records the outcome

Aetherfy decides whether a run succeeded from the process exit code alone.

Exit codeRun recorded as
0, or an unknown or unparseable codecompleted
Any known non-zero codefailed — the reason is kept with the run

An uncaught exception exits non-zero on every runtime Aetherfy supports, so a crash fails the run without you writing any error handling. You do not need to call sys.exit(1) or process.exit(1) yourself to mark a failure, though doing so is fine.

Some failures carry extra detail on the run record:

FailureHow Aetherfy records it
Out of memoryA failure, annotated (oom_killed) in the run’s error message
Host or infrastructure failure“the run’s machine failed (host/infrastructure failure)”
The run ended but its outcome never reached Aetherfy“the run ended without reporting its outcome”
No machine accepted the run within 2 minutes of Aetherfy starting to hand it over“The run was never accepted by a machine, so your code was not executed.”
The run was still waiting to be handed to a machine 7 minutes after it was created“The run was never handed to a machine, so your code was not executed.”
The run’s image rebuild had not started 7 minutes after the run was created“The run’s image rebuild never started, so your code was not executed.”
The run’s image rebuild had not finished 7 minutes after the run was created“The run’s image rebuild did not finish, so your code was not executed.”

Signal deaths map to the conventional 128+N codes:

SignalExit code
SIGKILL, including an out-of-memory kill137
SIGTERM143

Shutdown signals on Aetherfy tasks

When Aetherfy needs to stop a task, it sends SIGTERM to your process and every process it started, and gives your process a 10-second grace to exit before it sends SIGKILL to everything the run started. Use that window to flush a buffer or close a connection; do not use it for real work. The run is then reported, and the machine stops.

When a run ends — cleanly, with an error or killed — anything it started that is still running is killed before the run is reported, so each run on a reused machine starts clean.

The longer 120-second grace period applies to service agents, which may have requests in flight when they are asked to stop — it is described on /agents. A task has no such window: plan on 10 seconds.

Logging from an Aetherfy task

Everything your program writes to stdout and stderr becomes the run’s logs. Plain print in Python and console.log in Node are the supported logging interface on Aetherfy — there is no logging SDK to install and no special sink to configure.

print("starting nightly rollup", flush=True)
console.log('starting nightly rollup');

Aetherfy tags each record with the stream it came from (stdout, stderr, or system), so you can filter by stream when reading logs back. Retrieval flags, retention, and the per-line caps are on /agents/runs-and-logs.

Reading the input payload from Aetherfy

Input on Aetherfy arrives as a file. Before your entrypoint starts, the run’s payload is written to disk on the machine and its path is put in AETHERFY_SPAWN_PAYLOAD_PATH. Read that file and parse it as JSON. Nothing crosses the network for it.

The empty case is the normal case. A scheduled fire, or a manual run started without input, gets a file holding {}. Write the task so that no input is the path it takes most often, and treat any input as an optional override.

The payload is for parameters and references, not data. Aetherfy holds it to an inline cap of 256 KB: a spawn or a manual run above it is refused with 413 RUN_PAYLOAD_TOO_LARGE before any run is created. Anything larger belongs in a collection in your Aetherfy vector database, with its id in the payload.

Every run also receives these three variables:

VariableContains
AETHERFY_API_URLBase URL of the Aetherfy API
AETHERFY_SPAWN_IDThis run’s id
AETHERFY_API_KEYAn Aetherfy API key issued to this run, alongside AETHERFY_SPAWN_ID. It is set per run and stops working when the run ends, so read it from the environment each time rather than caching it beyond the run

They are what the fallback uses. If AETHERFY_SPAWN_PAYLOAD_PATH is unset, the machine could not write the file, and the same payload is available over HTTP — a fallback to keep in the code, not the path to build on:

GET {AETHERFY_API_URL}/deployments/{AETHERFY_SPAWN_ID}/payload Authorization: Bearer {AETHERFY_API_KEY}

The response shape is:

{"payload": {}}

A run started by a parent agent carries its payload through the same endpoint, so one implementation covers scheduled fires, manual runs, and spawned runs.

Returning a result from an Aetherfy task

Output on Aetherfy mirrors input. Before your entrypoint starts, Aetherfy puts the path of a file in AETHERFY_SPAWN_RESULT_PATH. Write JSON to that file and Aetherfy stores it on the run when your process ends. Nothing crosses the network for it, and there is no call to make.

Returning nothing is the normal case. Most tasks do their work and exit, and a run that writes no file is recorded as returning nothing — not as failing to return something.

The result is for answers and references, not data. Aetherfy holds it to the same inline cap as the payload, and anything larger belongs in a collection in your Aetherfy vector database, with its id in the result.

A result Aetherfy could not accept does not fail the run: the exit code is still the run’s outcome, and the reason is recorded beside the empty result.

What the run wroteWhat the run records
Nothingresult empty, result_error empty
JSON within the capresult holds it, result_error empty
More than the capresult empty, result_error is too_large
Something that is not JSONresult empty, result_error is not_json

Whoever started the run reads both fields back from the run itself — a parent agent, the CLI, or your own code. Aetherfy offers a plain read and a waiting read, and both are on /agents/api-deployments; the run history on /agents/api-runs says which runs answered without carrying the answers.

This is the path to use whenever a run has something to tell its caller, including a run in another region and a run started by a schedule. Code running inside one machine needs none of it — there a result is a return value.

Complete Python example

The SDK ships in the aetherfy-vectors distribution, which the standard runtime images preinstall — so on a plain agent this imports with nothing in your requirements.txt. A version you pin yourself wins over the preinstalled one.

"""An Aetherfy task: pulls its input payload, works, returns a result.""" from aetherfy_agent import payload, write_result def main(): data = payload() # A scheduled fire sends no input, so this is the normal path. target_date = data.get("date") if target_date is None: print("no input payload: processing the default window", flush=True) else: print(f"input payload requested date {target_date}", flush=True) # Small and by reference: an answer, not the data behind it. write_result({"rows": 128, "date": target_date}) print("done", flush=True) if __name__ == "__main__": main()

write_result is a task-only call: a result belongs to a run, and a service agent has no runs. It refuses a value over the inline cap rather than letting Aetherfy drop it silently, so you learn at the write that the answer belongs in a collection.

The same task without the SDK

Nothing above is a new protocol — the helper reads the same environment variables and the same files this page documents. Standard library only:

"""The payload and result contract, hand-rolled.""" import json import os import urllib.request def fetch_payload(): """Return this run's input payload, or {} when the run was given none.""" path = os.environ.get("AETHERFY_SPAWN_PAYLOAD_PATH") if path: with open(path, encoding="utf-8") as fh: return json.load(fh) or {} # Fallback: the machine could not write the file. Same bytes, over HTTP. api_url = os.environ["AETHERFY_API_URL"] spawn_id = os.environ["AETHERFY_SPAWN_ID"] api_key = os.environ["AETHERFY_API_KEY"] request = urllib.request.Request( f"{api_url}/deployments/{spawn_id}/payload", headers={ "Authorization": f"Bearer {api_key}", # Always set an explicit User-Agent. Python's urllib otherwise sends # "Python-urllib/3.x", a signature many CDNs' bot protection blocks # outright — producing a 403 that reads like an auth failure. "User-Agent": "nightly-report/1.0", }, ) with urllib.request.urlopen(request, timeout=30) as response: body = json.load(response) return body.get("payload") or {} def write_answer(value): """Return `value` to whoever started this run. A no-op when Aetherfy gave this machine no result path — the run still succeeds. The SDK raises there instead, because a library that discards the value it was handed is worse than one that says so.""" path = os.environ.get("AETHERFY_SPAWN_RESULT_PATH") if not path: return with open(path, "w", encoding="utf-8") as fh: json.dump(value, fh)

Complete Node example

The same helper on the aetherfy-vectors/agent subpath, preinstalled on the standard node20, node22 and bun images:

/** An Aetherfy task: reads its input payload, works, returns a result. */ const { payload, writeResult } = require('aetherfy-vectors/agent'); async function main() { const data = await payload(); // A scheduled fire sends no input, so this is the normal path. const targetDate = data.date; if (targetDate === undefined) { console.log('no input payload: processing the default window'); } else { console.log(`input payload requested date ${targetDate}`); } // Small and by reference: an answer, not the data behind it. await writeResult({ rows: 128, date: targetDate ?? null }); console.log('done'); } main().catch((error) => { console.error(error); process.exit(1); });

Hand-rolled, the same two calls are readFile on AETHERFY_SPAWN_PAYLOAD_PATH with the HTTP fallback above, and writeFile on AETHERFY_SPAWN_RESULT_PATH — the Python block just above shows both in full, and the contract is identical.

Environment variables Aetherfy provides

Beyond the three payload variables above, every Aetherfy agent machine carries:

VariableContains
AETHERFY_AGENT_IDThe agent’s id
AETHERFY_AGENT_NAMEThe agent’s name
AETHERFY_REGIONThe region this machine is running in
AETHERFY_WORKSPACEThe agent’s workspace — present only when the agent declares one. An agent with no workspace carries no such variable, and its clients stay workspaceless
AETHERFY_VECTORS_URLBase URL for the Aetherfy vector database
AETHERFY_SPAWN_URLA template for the spawn endpoint, .../agents/{id}/spawn with a literal {id}: substitute this agent’s own id (AETHERFY_AGENT_ID) before calling it. The SDK’s spawn helper does this for you
AETHERFY_DEPLOYMENT_IDThe deployment this machine is running. On a task agent this is the current run, set per run like AETHERFY_SPAWN_ID
AETHERFY_PARENT_AGENT_IDThe agent that spawned this run — present only on spawned runs
AETHERFY_WORKSPACE_AGENTSEvery agent in this workspace, this one included, comma-separated — present only when the agent declares a workspace
AETHERFY_MODEL_NAMEThe model name recorded on the agent — present only when one is set
AETHERFY_SPAWN_PAYLOAD_PATHPath of the file holding this run’s payload — see the payload section above
AETHERFY_SPAWN_RESULT_PATHPath of the file this run may write its result to — see the result section above
AETHERFY_RUN_INLINE_MAX_BYTESThe inline cap in bytes, the same number for a payload and a result. Read it rather than hardcoding 256 KB: it is the platform’s number and it can move
AETHERFY_MEMORY_MBThis machine’s memory, in MB — the memory_mb it was deployed with
AETHERFY_VCPUSThis machine’s shared vCPUs, derived from its memory — see /platform/limits. Size an in-machine pool from this

All of these are plain environment variables: os.environ["AETHERFY_VCPUS"] in Python, process.env.AETHERFY_VCPUS in Node. Nothing has to be imported to read them, and they are set before your entrypoint starts.

Spawning from code, and reading the answer back

AETHERFY_SPAWN_URL is the endpoint a run uses to start another task agent. The SDK’s spawn substitutes this agent’s own id into it, sends the request, and returns once the run is recorded; wait then holds one request open until that run finishes, instead of polling.

from aetherfy_agent import spawn, wait run = spawn("nightly-rollup", {"date": "2026-09-08"}) finished = wait(run.spawn_id, timeout_seconds=45) if finished.state == "completed": print(finished.result)
const { spawn, wait } = require('aetherfy-vectors/agent'); const run = await spawn('nightly-rollup', { date: '2026-09-08' }); const finished = await wait(run.spawn_id, 45); if (finished.state === 'completed') { console.log(finished.result); }

Acceptance is not execution. spawn returns when the run is recorded and its deploy queued. Aetherfy never queues a run behind another: a spawn aimed at an agent whose machines are all busy gets a machine of its own and runs at once (see the machine section below), and a spawn over the account’s runs-in-flight limit is refused with 429 AGENT_RUN_CONCURRENCY_LIMIT_EXCEEDED.

A wait that times out is not an error. It returns the run exactly as it stands, with state still active; read state and call again. The bound is 1 to 60 seconds, and waiting longer is a second call rather than a bigger number. result is the same read with no waiting.

Both calls are the REST endpoints on /agents/api-deployments, and hand-rolling them is a POST to the substituted spawn URL and a GET on the run — the helper adds no protocol.

The entire AETHERFY_ prefix is reserved on Aetherfy — you cannot set a secret that would shadow one of these. See /agents/secrets.

The machine a task run gets on Aetherfy

A task agent has one machine in each region it is deployed in — at rest. It can briefly have more, when runs overlap; see below. Between runs a machine is suspended, not destroyed; a run resumes it and starts your entrypoint as a fresh process. Do not rely on anything on the machine surviving between runs — a resume brings the previous state back, a cold start does not, and Aetherfy replaces the machine when it must. Anything you need to keep, write somewhere durable before you exit.

One machine runs one thing at a time, and every run passes the same gate. A run that finds the agent’s machines in its region all busy — whether you started it by hand, a schedule fired it, or a parent spawned it — gets a machine of its own from the same image, and both runs execute at once. The busy machine’s run is left to finish.

A run on a new machine waits for that machine to start. A machine resting between runs resumes and starts your entrypoint in about a second. An overlapping run’s new machine is a cold start, typically 15–20 seconds before your entrypoint begins — so “at once” means without waiting for the busy run, not without a start-up. The one limit is the account’s runs-in-flight limit: past it, a manual run or a spawn is refused with 429 AGENT_RUN_CONCURRENCY_LIMIT_EXCEEDED, and a scheduled occurrence is skipped and recorded with that same code. Each of those machines is an agent-region like any other and is billed like one — see Pricing . They are cleaned up for you once they have been idle a few minutes; one machine per region stays.

Aetherfy never queues a run behind another. A queued run would have no honest start time, and every run is at most once.

Runs of the same agent can overlap. Whatever triggers them — a schedule firing while the previous run is still going, a second afy run, several spawns — two or more runs of one agent can execute at the same time, up to your plan’s runs-in-flight limit. Aetherfy does not serialise them. A task that writes shared state must be idempotent (running it twice at once leaves the same result as once) or take its own lock — for example a row it claims in your database, or a key in your vector store — before it does work that must not happen twice.

How many runs you can have executing at once across the whole account is a plan limit — see the runs-in-flight row on Limits. It counts every task run, however it was triggered.

Two consequences follow for parallel work on Aetherfy:

You wantDo this
Many pieces of work in parallelFan out inside the machine — a thread pool for I/O-bound work such as model calls, a process pool for CPU-bound work. Size it from AETHERFY_VCPUS and AETHERFY_MEMORY_MB, and pick a larger memory_mb for more headroom: the 8192 step carries four shared vCPUs
Work that must run on separate machinesSpawn — including the same child several times over. Each run that finds every machine busy gets one of its own, so N spawns of one child give you N runs executing at once

Fanning out inside the machine is still the CHEAP kind of parallelism: no extra machines, no cold starts, and no extra awake time beyond the run itself. Reach for it first. Spawning scales out instead of up — a separate machine, its own memory, its own crash blast radius — and each of those machines is billed for the time it is awake, so a fan-out of fifty spawns costs about fifty times what one run costs. Use it when the work genuinely needs isolation, a different memory size, or more capacity than one machine has; use in-machine fan-out when it does not. A spawned child is a task agent like any other, priced by its memory, per region. The account-wide ceiling on runs executing at once is the runs-in-flight row on Limits.

Spawning also has a latency shape worth knowing. Every spawn is recorded by the Aetherfy control plane before the child is resumed, and the control plane runs in one region, so a spawn from another region pays a round trip there and back before your child’s code starts. That is nothing against a run that lasts seconds or minutes. For work that needs sub-second fan-out, fan out inside the machine, which touches no control plane at all; for a direct conversation between agents, use a service agent and call it at its URL, which every agent in its workspace receives as AETHERFY_AGENT_<NAME>_URL (see Workspaces). Keep the payload small — it is for parameters, not data — and pass anything large by reference to a collection in your Aetherfy vector database, which every agent in the region reaches over the private network.

Fan-out inside the machine, in Python

The shape of a dynamic fan-out on Aetherfy: invent the work at runtime, hand each piece to an anonymous worker that shares the process and everything in it, and get the results back in-band. call_your_model stands for whichever model client you use.

"""An Aetherfy task that fans out inside its own machine.""" from aetherfy_agent import fan_out, machine, payload def invent_tasks(data): """Decide the work at runtime — here, one prompt per item in the input.""" return [f"Summarise the change in {item}" for item in data.get("items", [])] def worker(prompt): """One anonymous worker: takes a prompt, returns its answer. A model call is I/O-bound, so a thread is the right worker for it.""" return call_your_model(prompt) def main(): data = payload() # Results come back in INPUT order, and no failure is swallowed: the # lowest-indexed exception is re-raised once every worker has finished. # Width defaults to vcpus x 8 for threads, which is the pool for waiting. results = fan_out(worker, invent_tasks(data)) for result in results: print(result, flush=True) if __name__ == "__main__": main()

CPU-bound work wants processes, and the default width follows the pool — threads get vcpus * 8 because they mostly wait, processes get vcpus, one per core:

from aetherfy_agent import fan_out digests = fan_out(hash_one, files, kind="processes")

fan_out prints one line before it starts, so how wide a run went is visible in its logs afterwards:

aetherfy: fanning out 32 wide on 4 vCPU / 8192 MB (120 tasks)

machine() is where that width comes from, and you can size your own pool from it: vcpus, memory_mb and region, as whole numbers rather than the strings the environment carries. Memory is the limit that ends the whole run if you cross it — an out-of-memory kill is exit 137 above.

Fan-out inside the machine, in Node

The same shape on the Node runtimes. fanOut is promise concurrency, which is all I/O-bound work needs; move CPU-bound work to worker_threads, no wider than machine().vcpus.

// An Aetherfy task that fans out inside its own machine. const { fanOut, payload } = require('aetherfy-vectors/agent'); function inventTasks(data) { // Decide the work at runtime — here, one prompt per item in the input. return (data.items ?? []).map((item) => `Summarise the change in ${item}`); } async function worker(prompt) { // One anonymous worker: takes a prompt, returns its answer. return callYourModel(prompt); } async function main() { const data = await payload(); // Input order, nothing swallowed, and the same line printed to the logs. const results = await fanOut(worker, inventTasks(data)); for (const result of results) console.log(result); } main().catch((err) => { console.error(err); process.exit(1); });

Without the SDK

Both of those are a pool and a print. Hand-rolled in Python it is a ThreadPoolExecutor sized from AETHERFY_VCPUS, and in Node a slice loop over Promise.all — the environment variables in the table below are the whole input. The helper exists so the sizing, the input-order guarantee and the log line stop being retyped per task, not because the platform requires it.

The Aetherfy runtimes are Python, Node and Bun, and the pattern is the same on each: invent the work, fan it out in-process, size the width from the two variables. A custom container (runtime: dockerfile) is not covered by this page.

A deploy that lands while a run is in flight does not interrupt it. The run finishes on the image it started with, and the next run uses the new image, so you never have to time a deploy around a schedule.

At-most-once execution on Aetherfy

A scheduled occurrence on Aetherfy fires at most once. A failed run is not retried; the next attempt is simply the next scheduled occurrence. Aetherfy has no automatic retry, no backoff, and no dead-letter queue for runs.

Two consequences are worth designing around explicitly.

ConsequenceWhat to do about it
A run that failed part-way still left its earlier effects behind, and you may re-run the same period by handMake the work idempotent — upsert rather than insert, and key writes by the period they cover so a second attempt overwrites rather than duplicates
An occurrence can fail to happen at allDo not assume every occurrence ran. Derive the window you process from your own data watermark (“everything since the last row I wrote”) rather than from the clock (“the last 24 hours”)

Written that way, a task that misses a night catches up automatically on the next fire, because the window it computes is wider. Written the other way, the missing night stays missing forever.

If you need every single occurrence processed with a guarantee, a schedule is the wrong tool on Aetherfy — put the work in a durable queue and let a task drain it.

The run route for an Aetherfy service

A type: service agent can be run too — on its schedule, by hand with afy run, or by a parent’s spawn. A service has no script to execute, so a run of a service is one HTTP request that Aetherfy sends to a route on your own server:

POST /aetherfy/run X-Aetherfy-Run: <run id> Content-Type: application/json

The body names the run, what started it, and its input:

{"run_id": "3c9a5f10-7b2e-4d81-a6f4-19e0c8b7d224", "trigger": "cron", "payload": {}}

trigger is cron for a scheduled run and manual for afy run. payload is {} when the run has no input, never null, and it is held to the same 256 KB inline cap as a task’s.

Your response decides the run, the way a task’s exit code does:

Your routeRun recorded as
Answers with any 2xxcompleted — a JSON body is the run’s result
Answers with anything elsefailed — “the service answered HTTP 500 at POST /aetherfy/run”, with your status
Does not answer within the run timeoutfailed — “the service did not answer POST /aetherfy/run within the run timeout”
Cannot be reachedfailed — “the service could not be reached at POST /aetherfy/run”

The run timeout is the same 60 minutes as the task backstop below. Your service is not stopped when a run times out; only the run is recorded as failed.

A run that wakes a service whose process has to start again — a cold wake rather than a resume — waits for your server. Until it is listening, Aetherfy asks again, backing off, for up to the same 2 minutes any run has to be accepted. A server that starts in that time takes the run; one that does not leaves the run failed with “The run was never accepted by a machine, so your code was not executed.”

A run that is still waiting on your route when the service itself goes away is failed at that moment, with a reason naming what happened, rather than waiting out the timeout:

What happened to the serviceRun recorded as
A new version was deployed, or an older one rolled back tofailed — “a new version of the service was deployed while the run was in flight”
You stopped itfailed — “the service was stopped while the run was in flight”
The account reached its spend limitfailed — “the service was paused at the account’s spend limit while the run was in flight”
The account was suspendedfailed — “the service was stopped because the account was suspended while the run was in flight”
It was archivedfailed — “the service was archived while the run was in flight”
Its app is no longer on the compute planefailed — app_lost_on_provider
Its image is no longer available on the compute planefailed — image_lost_on_provider

The last two reasons are the agent’s own failure codes for the same loss, described in the agent lifecycle API; a redeploy recovers the agent.

A run’s duration in the run history is measured from the moment the service’s machine accepted the run to the moment its outcome was reported.

A 2xx body is read with the rules a task’s result file follows: JSON within the cap is stored as the result, more than the cap is recorded as too_large, anything that is not JSON as not_json. An empty body — a 204, or a 200 with nothing in it — returns nothing. Redirects are not followed: a 3xx is an answer, and it fails the run.

A minimal Python route, in a FastAPI service that exports app:

from fastapi import FastAPI, Request app = FastAPI() @app.post("/aetherfy/run") async def run(request: Request): body = await request.json() rows = rebuild_index(since=body["payload"].get("since")) return {"rows": rows}

Four things differ from a task, each on purpose:

DifferenceWhat it means for your service
Runs are concurrentA service already answers requests concurrently, so a run that starts while the previous one is still going is simply another request. A service run adds no machine, so it does not count toward the runs-in-flight limit: a service’s scheduled run is never skipped, and a manual run or a spawn of a service is never refused by that limit
Only Aetherfy reaches the routeA request to /aetherfy/run from outside is answered 404 before it reaches your server, so nobody else can run your service by guessing the path
Your logs stay your service’sWhat your route prints is part of the service’s own logs, because concurrent runs cannot honestly be told apart there. afy logs <name> --run <id> shows the lines Aetherfy writes about that run: when it was sent, what your route answered, and how long it took
The API key is the service’sA run of a service is sent no key of its own. Your server’s environment is fixed when it starts, so it keeps authenticating with the AETHERFY_API_KEY its deployment was given

A service run does not count toward the runs-in-flight limit on your plan: that limit bounds machines, and a service run adds none. For the same reason it is not a task run in your usage figures: the time it took is your service machine’s own awake time, which is already counted as the service’s.

Where a spawned service run goes

A run you start by hand or on a schedule goes to the service’s own region. A run a parent spawns goes to a region where the service has a running machine, preferring the parent’s region: if the service runs there, the request stays in that region; if it does not, Aetherfy sends it to another region of the service’s live release. Only when the service has no running machine anywhere is the spawn refused, with 409 AGENT_SERVICE_NOT_RUNNING naming the regions its release is in. A spawn never starts or deploys a service.

The 60-minute backstop on Aetherfy

A run still executing after 60 minutes is terminated by Aetherfy and recorded as failed. Aetherfy checks every 5 minutes, so a run past the limit ends within 5 minutes of reaching it. The machine was awake for that time, so it counts toward the agent’s uptime the same way any awake time does — see Billing and spend caps.

This is a cost-safety backstop, not a target, and it is not configurable. It exists so a task stuck in a loop cannot bill indefinitely.

If your work legitimately exceeds an hour, split it into smaller scheduled batches — for example, process one hour of data per fire on an hourly schedule instead of a day of data per fire on a daily one. Each fire then finishes well inside the backstop, and the watermark pattern above makes the batches self-correcting.

Last updated on