Ship AI features agents that act
One API for models, retrieval, and tools, with evals and tracing from the first request.
- Rated4.9/5
- SOC 2Type II
- Uptime99.99%
support-agent
RunningDoneWhere is order 4821, and can I still change the delivery address?
Order 4821 left the warehouse this morning and should arrive Thursday. You can still change the address until it reaches the local depot, so I opened the change form for you.
Trace
- search_docsshipping-policy.md, 3 chunks84 ms
- get_order4821 · in transit120 ms
- check_address_changeAllowed until the local depot96 ms
- Grounded0.97
- Latency1.2 s
- Cost$0.004
Everything between the prompt and production
Agents, retrieval, routing, and observability that share one workspace and one bill.
Agents with real tools
Write tools as typed functions. The agent plans, calls them, and retries failed steps, with every call logged.
support-agent.ts- 1@tool({ retries: 2 })
- 2export async function lookupOrder(id) {
- 3 const order = await orders.get(id)
- 4 await audit.log('lookup', id)
- 5 return order.status
- 6}
Retrieval over your own data
Sync documents, tickets, and tables into a managed vector index. Answers cite the chunks they came from.
docs-index
SyncedDocuments
Queries
Recall
Latency
Model routing and fallbacks
Send each request to the right model by cost, latency, or task, and fail over when a provider slows down.
Inference pool
- 1 GPU · 24 GB
- 4 GPU · 96 GB
- 8 GPU · 320 GB
- 1 GPU · 24 GB
Traces and evals
Follow every prompt, tool call, and token. Run eval suites on each release and catch regressions before users do.
Traces- 10:12:00agent.run support-agent 1.24s
- 10:13:01tool: lookupOrder 84ms
- 10:14:02retrieve: docs-index 6 chunks
- 10:15:03llm: completion 412 tokens
- 10:16:04eval: groundedness 0.94 pass
- 10:17:05guardrail: pii redacted 2 fields
- 10:18:06agent.run triage-agent 0.88s
- 10:19:07llm: fallback to secondary model
- 10:12:08agent.run support-agent 1.24s
- 10:13:09tool: lookupOrder 84ms
- 10:14:00retrieve: docs-index 6 chunks
- 10:15:01llm: completion 412 tokens
- 10:16:02eval: groundedness 0.94 pass
- 10:17:03guardrail: pii redacted 2 fields
- 10:18:04agent.run triage-agent 0.88s
- 10:19:05llm: fallback to secondary model
Tokens / min
From first prompt to production in three steps
Pick a model, ground it in your data, and ship with guardrails, evals, and tracing turned on.
- Chat
- Embeddings
- Vision
- Speech
- Rerank
- Fine-tune
- 01Connected to docs-bucket
- 02Found 1,240 documents
- 03Chunked into 18,600 passages
- 04Embeddings created in 3m 12s
- 05Index docs-index is ready
- 06Sync scheduled every 15 minutes
- PII redaction
- Per-key rate limits
- Spend caps and alerts
- Eval suite on each release
- Full request tracing
- Automatic model fallback
Customers
AI features your users can trust
Product and platform teams run agents, search, and copilots in production on the platform.
“Our support agent resolves a third of tickets on its own. Traces showed us exactly where it went wrong before we let it answer customers.”
“Search over our docs went from a hackathon demo to production in two weeks. Every answer links to its source, so people trust it.”
“The eval suite runs on each pull request. We catch prompt regressions in review instead of in support tickets.”
“We swapped models twice in one quarter without touching product code. Routing rules and evals made each switch a config change.”
“Spend caps per key let us give every team access to models without a surprise bill at the end of the month.”
“Fallbacks kept our shopping assistant online through two provider outages. Customers never noticed.”
“Our support agent resolves a third of tickets on its own. Traces showed us exactly where it went wrong before we let it answer customers.”
Head of Support Engineering, Northstar Systems “We swapped models twice in one quarter without touching product code. Routing rules and evals made each switch a config change.”
CTO, Brightlane Labs “Search over our docs went from a hackathon demo to production in two weeks. Every answer links to its source, so people trust it.”
Platform Lead, VectorGrid “Spend caps per key let us give every team access to models without a surprise bill at the end of the month.”
Founder, Crestfield Digital “The eval suite runs on each pull request. We catch prompt regressions in review instead of in support tickets.”
VP of Product, Summit Cloud “Fallbacks kept our shopping assistant online through two provider outages. Customers never noticed.”
Engineering Manager, Orbit Commerce
Pricing
Pay for the tokens you use
Start free, then pay per token with a flat platform fee. No seat licenses and no minimum commit.
Estimate your monthly bill
- Input tokens
$0.50 / M tokensFirst 1 M tokens free
- Output tokens
$1.50 / M tokens
- Embeddings
$0.05 / M tokensFirst 10 M tokens free
- Vector storage
$0.25 / GBFirst 1 GB free
Estimated monthly cost
$/ month
- Pro · base fee
- $20.00
- Input tokens
- $99.50
- Output tokens
- $90.00
- Embeddings
- $24.50
- Vector storage
- $4.75
Estimates exclude taxes. Tokens are metered per request and billed monthly.
- Model routing and fallbacks
- Evals and tracing
- Spend caps per key
- Email support
FAQ
Questions, answered
Can't find what you're looking for? Our team usually replies within a few hours.
Contact support01Which models can I use?
Chat, embedding, vision, and speech models from several providers, plus open models we host. All of them use the same request format.
02Is my data used to train models?
No. Prompts, completions, and indexed documents are never used for training, and you choose how long logs are kept.
03How are tokens counted?
Input and output tokens are counted per request and shown in each trace. Usage rolls up per key, project, and workspace.
04Can I cap what a team spends?
Yes. Set a monthly cap and alert thresholds on each key or project. Requests stop or fall back to a cheaper model when a cap is reached.
05How do evals work?
Write test cases with expected behavior, then run the suite from the CLI or on each pull request. Scores are stored next to the traces.
06Can I run it in my own cloud?
Enterprise plans include private deployments in your own account and region, with the same API and dashboard.