Ship AI features agents that act

One API for models, retrieval, and tools, with evals and tracing from the first request.

support-agent

RunningDone

Where is order 4821, and can I still change the delivery address?

Order 4821 left the warehouse this morning and should arrive Thursday. You can still change the address until it reaches the local depot, so I opened the change form for you.

Trace

  1. search_docsshipping-policy.md, 3 chunks84 ms
  2. get_order4821 · in transit120 ms
  3. check_address_changeAllowed until the local depot96 ms
  • Grounded0.97
  • Latency1.2 s
  • Cost$0.004

Powering AI features at product teams of every size

Everything between the prompt and production

Agents, retrieval, routing, and observability that share one workspace and one bill.

  • Agents with real tools

    Write tools as typed functions. The agent plans, calls them, and retries failed steps, with every call logged.

    support-agent.ts
    1. 1@tool({ retries: 2 })
    2. 2export async function lookupOrder(id) {
    3. 3 const order = await orders.get(id)
    4. 4 await audit.log('lookup', id)
    5. 5 return order.status
    6. 6}
  • Retrieval over your own data

    Sync documents, tickets, and tables into a managed vector index. Answers cite the chunks they came from.

    docs-index

    Synced
    • Documents

    • Queries

    • Recall

    • Latency

  • Model routing and fallbacks

    Send each request to the right model by cost, latency, or task, and fail over when a provider slows down.

    Inference pool

    • 1 GPU · 24 GB
    • 4 GPU · 96 GB
    • 8 GPU · 320 GB
    • 1 GPU · 24 GB
  • Traces and evals

    Follow every prompt, tool call, and token. Run eval suites on each release and catch regressions before users do.

    Traces
    • 10:12:00agent.run support-agent 1.24s
    • 10:13:01tool: lookupOrder 84ms
    • 10:14:02retrieve: docs-index 6 chunks
    • 10:15:03llm: completion 412 tokens
    • 10:16:04eval: groundedness 0.94 pass
    • 10:17:05guardrail: pii redacted 2 fields
    • 10:18:06agent.run triage-agent 0.88s
    • 10:19:07llm: fallback to secondary model
    • 10:12:08agent.run support-agent 1.24s
    • 10:13:09tool: lookupOrder 84ms
    • 10:14:00retrieve: docs-index 6 chunks
    • 10:15:01llm: completion 412 tokens
    • 10:16:02eval: groundedness 0.94 pass
    • 10:17:03guardrail: pii redacted 2 fields
    • 10:18:04agent.run triage-agent 0.88s
    • 10:19:05llm: fallback to secondary model

    Tokens / min

From first prompt to production in three steps

Pick a model, ground it in your data, and ship with guardrails, evals, and tracing turned on.

  • Chat
  • Embeddings
  • Vision
  • Speech
  • Rerank
  • Fine-tune

Customers

AI features your users can trust

Product and platform teams run agents, search, and copilots in production on the platform.

“Our support agent resolves a third of tickets on its own. Traces showed us exactly where it went wrong before we let it answer customers.”
Head of Support Engineering, Northstar Systems
“Search over our docs went from a hackathon demo to production in two weeks. Every answer links to its source, so people trust it.”
Platform Lead, VectorGrid
“The eval suite runs on each pull request. We catch prompt regressions in review instead of in support tickets.”
VP of Product, Summit Cloud
“We swapped models twice in one quarter without touching product code. Routing rules and evals made each switch a config change.”
CTO, Brightlane Labs
“Spend caps per key let us give every team access to models without a surprise bill at the end of the month.”
Founder, Crestfield Digital
“Fallbacks kept our shopping assistant online through two provider outages. Customers never noticed.”
Engineering Manager, Orbit Commerce
Read customer stories

Pricing

Pay for the tokens you use

Start free, then pay per token with a flat platform fee. No seat licenses and no minimum commit.

Estimate your monthly bill

  • Input tokens

    $0.50 / M tokensFirst 1 M tokens free

  • Output tokens

    $1.50 / M tokens

  • Embeddings

    $0.05 / M tokensFirst 10 M tokens free

  • Vector storage

    $0.25 / GBFirst 1 GB free

FAQ

Questions, answered

Can't find what you're looking for? Our team usually replies within a few hours.

Contact support
01Which models can I use?

Chat, embedding, vision, and speech models from several providers, plus open models we host. All of them use the same request format.

02Is my data used to train models?

No. Prompts, completions, and indexed documents are never used for training, and you choose how long logs are kept.

03How are tokens counted?

Input and output tokens are counted per request and shown in each trace. Usage rolls up per key, project, and workspace.

04Can I cap what a team spends?

Yes. Set a monthly cap and alert thresholds on each key or project. Requests stop or fall back to a cheaper model when a cap is reached.

05How do evals work?

Write test cases with expected behavior, then run the suite from the CLI or on each pull request. Scores are stored next to the traces.

06Can I run it in my own cloud?

Enterprise plans include private deployments in your own account and region, with the same API and dashboard.

Free tier · No credit card

Build your first agent today

Get an API key and make your first request in under five minutes.

$ cloud agents deploy support-bot

  1. Build21 s
  2. Evals24/24
  3. Deploy9 s

support-bot.example.com

Buy NowTheme Details