How a team of four at Summit Cloud runs AI workloads at scale
An AI startup runs inference APIs and background jobs for millions of requests a day without a dedicated ops team.
Results
A small team, a large workload
- Jobs a day
- M
- Embedding, summarizing, and indexing documents.
- Engineers
- Running the whole platform, on-call included.
- Less ops time
- %
- Compared with their first self-managed setup.
The challenge
Summit Cloud turns large document sets into searchable knowledge bases. Its first version ran on self-managed clusters, and the founding engineers spent more time on queues, scaling scripts, and certificates than on the product.
As customers grew, the workload became spiky: a single upload could add hundreds of thousands of background jobs in minutes.
The solution
The team split the product into a web service for the API and workers for background jobs, both defined in one config file and deployed from the same repository. Workers scale on queue depth, and the API scales on request latency.
The CLI and API let them script new customer environments, and private networking keeps the vector store reachable only from their own services.
The results
Four engineers now run the whole platform, including on-call. Large uploads finish in minutes instead of hours, and the time they used to spend on infrastructure goes into the product.
“Running our AI workers here let a team of four handle the workload we expected to need fifteen people for.”