For AI apps
Ship AI apps without the GPU ops
Run inference APIs, agents, and long background jobs on one platform, with autoscaling that follows bursty model traffic.
- 1@task(retries=3, timeout='15m')
- 2async def answer(question):
- 3 docs = await search(question)
- 4 reply = await llm.run(question, docs)
- 5 await save_trace(question, reply)
- 6 return reply
AI teams in production
From prototype notebook to production agent
The hard part of AI apps is everything around the model call. That part is built in.
Jobs that outlast a request
Embedding, summarizing, and agent loops run as background jobs with retries and checkpoints.
Traffic that arrives in bursts
Inference APIs scale out on queue depth or CPU, then back down when the launch spike is over.
Vectors next to your data
Managed Postgres with vector search keeps embeddings, users, and metadata in one place.
Keys that stay private
Model provider keys live in encrypted environment variables, and workers talk over a private network.
Code
An agent and its worker in one config
Declare the API, the background worker, and the database together. Every push deploys all three.
- Streaming responses from web services
- Retries and timeouts per task
- Traces stored next to your data
1
services:
2
- type: web
3
name: agent-api
4
runtime: python
5
startCommand: uvicorn app:api
6
autoscaling:
7
minInstances: 1
8
maxInstances: 12
9
targetCPUPercent: 65
10
- type: worker
11
name: agent-worker
12
runtime: python
13
startCommand: python -m worker
14
databases:
15
- name: agent-db
16
extensions: [vector]
1
from cloud.tasks import task
2
3
@task(retries=3, timeout="15m")
4
async def index_document(doc_id: str):
5
doc = await fetch_document(doc_id)
6
chunks = split(doc.text, size=800)
7
vectors = await embed(chunks)
8
await db.save_vectors(doc_id, vectors)
1
$ cloud deploy
2
✓ agent-api live · 2 instances
3
✓ agent-worker live · 1 instance
4
✓ agent-db ready · vector enabled
Templates
Start from a working AI stack
Fork a starter with the worker, database, and API already wired together.
- BackendFastAPI serviceTyped Python API with background jobs and migrations.PythonRedisDeploy
- AIAI chat agentStreaming chat agent with tool calls and durable memory.PythonWorkflowsDeploy
- BackendGo microserviceLightweight HTTP service with private networking.GogRPCDeploy
- AIRAG pipelineDocument ingestion worker with vector search.WorkersPostgresDeploy
Customer story
Four engineers, a million jobs a day
An AI startup runs inference APIs and background jobs without a dedicated ops team.
“Running our AI workers here let a team of four handle the workload we expected to need fifteen people for.”
- Jobs a day
- M
- Embedding, summarizing, and indexing documents.
- Engineers
- Running the whole platform, on-call included.
FAQ
AI apps on the platform
What teams ask before they move an AI product into production.
Talk to an engineer01Do you run the models?
No. Your services call the model provider you choose, or a model you host in your own service. The platform runs everything around the call: APIs, workers, storage, and networking.
02How long can a background job run?
Each task sets its own timeout, up to several hours. Checkpoints let long jobs resume after a deploy or a failed step instead of starting over.
03Can I stream responses to the browser?
Yes. Web services support streaming responses and long-lived connections, so tokens reach users as soon as the model returns them.