We run the models
Operated open-weight inference on managed compute. No GPU ownership, no infrastructure to provision — you write the app, we run the models.
GPU-hour billing. Start experimenting immediately.
AI Applications Platform
Bespin Works runs, governs, and budgets every enterprise AI application — your keys or our compute, one endpoint, one audit trail.
Switch between operated compute and your own provider keys without changing the endpoint.
Operated open-weight inference on managed compute. No GPU ownership, no infrastructure to provision — you write the app, we run the models.
GPU-hour billing. Start experimenting immediately.
Provider pass-through. Route through your existing OpenAI, Anthropic, Azure, or Google keys — bring your own keys, we handle the rest.
Bespin Works governs, routes, and audits — never touches, resells, or marks up your provider spend. No vendor lock-in.
Unlike token-based gateways, Bespin Works provides a predictable cost-envelope and a signable governance model.
| Feature | Bespin Works | Typical Token Gateways |
|---|---|---|
| Billing Model | ✓ GPU-hour & Per-Project | Token Markup |
| Budget Control | ✓ Hard Monthly Caps | Post-payment Billing |
| Governance | ✓ Signable Summary | Self-serve Dashboard |
| Provider Lock-in | ✓ None (Neutral) | Implicit (Provider-specific) |
Hard caps, not surprise bills. Every project gets a monthly ceiling — hit it and we throttle, not bill.
Platform Preview
GPU-hour pricing means you know the cost before you deploy. One audit trail tracks every request to the user who made it. You build. We govern.
Platform Preview
Start as an experiment, graduate to production — on the same endpoint. No rebuilding, no re-platforming, no new infrastructure.
Conceptual preview
The short answers to the questions we hear most.
An AI Applications Platform is the layer that runs, governs, and budgets enterprise AI applications. SaaS gives you the application runtime; GPU clouds give you the inference substrate; gateways route between models. Bespin Works combines all three, and prices by the project, not the token.
Token-based gateways bill by the token, which makes invoices hard to predict. Bespin Works bills by GPU-hour and per project, with a hard monthly cap on every project, so you can model the cost before you deploy.
Every project runs under a per-project plan with a hard monthly cap. Operated projects bill by GPU-hour; provider projects bill a flat per-project fee. Hit the cap and we throttle, not bill.
No. In operated mode, Bespin Works runs open-weight models on managed compute. You write the application, we run the models. No GPU ownership, no infrastructure to provision.
Yes. In provider mode, you bring your own keys and Bespin Works routes, governs, and audits the traffic without touching, reselling, or marking up your provider spend.
You're not tied to a single model provider. Switch a project between operated compute and your own provider keys anytime, on the same endpoint, without re-platforming.
Every project gets a hard monthly cap, every request is attributed to a user, and every month you get a signable governance summary. Governance covers cost, access, and routing: one point of policy for all your enterprise AI.
Talk to us. We'll discuss your AI application roadmap, understand your use case, and set up a fit-check. Email ceo@bespin.works.
See how Bespin Works runs, governs, and budgets enterprise AI applications for your organization.