Scale

Autoscaling

Add and remove instances automatically as traffic, CPU, and memory change.

web-frontend

  • 2 instances · 1 CPU
  • 6 instances · 2 CPU
  • 18 instances · 4 CPU
  • 2 instances · 1 CPU

Capacity that follows demand

Set the limits once. The platform handles every spike and quiet hour in between.

  • CPU and memory targets

    Scale on average CPU, memory, or both, with thresholds you choose per service.

  • Min and max instances

    Keep a warm floor for fast responses and a ceiling that protects your budget.

  • Scale-down cooldowns

    Avoid flapping with cooldown windows that wait for traffic to settle.

  • Load-balanced by default

    New instances join the pool as soon as they pass health checks.

How autoscaling works

Three checks run every few seconds, so capacity is always one step ahead of traffic.

  • Min 2
  • Max 20
  • CPU 70%
  • Memory 80%

Configuration

Autoscaling in a few lines

Turn it on in the dashboard, or keep it in your config file next to the rest of your stack.

  • Per-service limits and targets
  • Change limits without a redeploy
  • Same config for staging and production
          
            
                1
                services:
              
                2
                  - type: web
              
                3
                    name: web-frontend
              
                4
                    runtime: node
              
                5
                    autoscaling:
              
                6
                      minInstances: 2
              
                7
                      maxInstances: 20
              
                8
                      targetCPUPercent: 70
              
                9
                      targetMemoryPercent: 80
              
          
        

When to use autoscaling

Autoscaling fits any service with traffic that changes over the day: storefronts, public APIs, webhooks, and AI endpoints that see bursts after a launch or a campaign. You pay for the instances you run, so a quiet night costs less than a busy afternoon.

What scales and what stays fixed

Web services and private services scale horizontally. Each new instance runs the same build and joins the same private network, so there is nothing to reconfigure. Background workers can scale on queue depth, and databases scale separately with their own plans.

  • Horizontal scaling for web and private services
  • Queue-based scaling for background workers
  • Independent limits for every environment

Keeping costs predictable

Your maximum instance count is a hard ceiling. When a service reaches it, you get an alert and the service keeps running at that size. Usage is billed per second, and the usage page shows exactly when each instance started and stopped.

Free tier · No credit card

Push code today. Be live before your coffee cools.

Connect a repository, pick a region, and get a production URL with HTTPS, logs, and autoscaling already switched on.

$ git push origin main

  1. Build34 s
  2. Deploy12 s
  3. Health checks3 s

your-app.example.com

Buy NowTheme Details