For AI apps

Ship AI apps without the GPU ops

Run inference APIs, agents, and long background jobs on one platform, with autoscaling that follows bursty model traffic.

agent.py
  1. 1@task(retries=3, timeout='15m')
  2. 2async def answer(question):
  3. 3 docs = await search(question)
  4. 4 reply = await llm.run(question, docs)
  5. 5 await save_trace(question, reply)
  6. 6 return reply

AI teams in production

From prototype notebook to production agent

The hard part of AI apps is everything around the model call. That part is built in.

  • Jobs that outlast a request

    Embedding, summarizing, and agent loops run as background jobs with retries and checkpoints.

  • Traffic that arrives in bursts

    Inference APIs scale out on queue depth or CPU, then back down when the launch spike is over.

  • Vectors next to your data

    Managed Postgres with vector search keeps embeddings, users, and metadata in one place.

  • Keys that stay private

    Model provider keys live in encrypted environment variables, and workers talk over a private network.

Code

An agent and its worker in one config

Declare the API, the background worker, and the database together. Every push deploys all three.

  • Streaming responses from web services
  • Retries and timeouts per task
  • Traces stored next to your data
AI quickstart
          
            
                1
                services:
              
                2
                  - type: web
              
                3
                    name: agent-api
              
                4
                    runtime: python
              
                5
                    startCommand: uvicorn app:api
              
                6
                    autoscaling:
              
                7
                      minInstances: 1
              
                8
                      maxInstances: 12
              
                9
                      targetCPUPercent: 65
              
                10
                  - type: worker
              
                11
                    name: agent-worker
              
                12
                    runtime: python
              
                13
                    startCommand: python -m worker
              
                14
                databases:
              
                15
                  - name: agent-db
              
                16
                    extensions: [vector]
              
          
        

Customer story

Four engineers, a million jobs a day

An AI startup runs inference APIs and background jobs without a dedicated ops team.

“Running our AI workers here let a team of four handle the workload we expected to need fifteen people for.”
Sofia LindqvistVP of Product, Summit CloudRead the story
Jobs a day
M
Embedding, summarizing, and indexing documents.
Engineers
Running the whole platform, on-call included.
Read customer stories

FAQ

AI apps on the platform

What teams ask before they move an AI product into production.

Talk to an engineer
01Do you run the models?

No. Your services call the model provider you choose, or a model you host in your own service. The platform runs everything around the call: APIs, workers, storage, and networking.

02How long can a background job run?

Each task sets its own timeout, up to several hours. Checkpoints let long jobs resume after a deploy or a failed step instead of starting over.

03Can I stream responses to the browser?

Yes. Web services support streaming responses and long-lived connections, so tokens reach users as soon as the model returns them.

Get started

Deploy your first agent today

Install the CLI, link a repository, and your API and worker go live together.

Install the CLI
npm install -g @cloud/cli
Buy NowTheme Details