Skip to main content
Coming soon
Dedicated inference

Make your model work without doing all the work.

We’re building a simpler way to run open-weight models on compute you control, served on private URLs. We’ll handle provisioning, runtime setup, and observability for you.

Request early access

You’ll need an ngrok.ai account to finish the request. We’d love to hear what you’re building! Or skip the questions. That’s cool, too.

Why run your own models, anyway?

When inference becomes a bigger part of your product, you may want more say in what runs, what it costs, and where it lives.

  • Match the model to the work. Classification, tool selection, and repeated analysis can be good fits for smaller, specialized models.

  • Make the economics fit. At high volumes, depending on your workload, hosting a model could lower your cost per task.

  • Choose where inference runs. Deploy in your own cloud account, with a supported compute provider you choose.

A model is only the beginning.

Getting a model running is one thing. Keeping it flippin’ fast, secure, and observable is another service for your team to maintain.

or, you could…

  1. Choose your model

    Choose an open-weight model that fits your workload and your budget.

  2. Choose your compute

    Pick a provider, region, and GPU. We provision the compute and configure the runtime for you.

  3. Send your first request

    Call your model through the AI Gateway. We’ve locked it down and layered in observability.

Your model on any compute you choose.

  • RunPod
  • DigitalOcean
  • Linode
  • Vultr
  • Verda
  • Scaleway
  • Hyperstack
  • Lambda

Start with one of these providers, or tell us who’s missing. We’re shaping it all up right now.