AI Applications Platform

Move fast on AI, and keep it under control.

Bespin Works runs, governs, and budgets every enterprise AI application — your keys or our compute, one endpoint, one audit trail.

Hard monthly caps One endpoint GPU-hour billing

Two channels. One platform.

Switch between operated compute and your own provider keys without changing the endpoint.

Operated

We run the models

Operated open-weight inference on managed compute. No GPU ownership, no infrastructure to provision — you write the app, we run the models.

GPU-hour billing. Start experimenting immediately.

Provider

Use your existing API keys

Provider pass-through. Route through your existing OpenAI, Anthropic, Azure, or Google keys — bring your own keys, we handle the rest.

Bespin Works governs, routes, and audits — never touches, resells, or marks up your provider spend. No vendor lock-in.

Engineered for Enterprise

Unlike token-based gateways, Bespin Works provides a predictable cost-envelope and a signable governance model.

Feature Bespin Works Typical Token Gateways
Billing Model ✓ GPU-hour & Per-Project Token Markup
Budget Control ✓ Hard Monthly Caps Post-payment Billing
Governance ✓ Signable Summary Self-serve Dashboard
Provider Lock-in ✓ None (Neutral) Implicit (Provider-specific)

Governance by Default

Hard caps, not surprise bills. Every project gets a monthly ceiling — hit it and we throttle, not bill.

Platform Preview

Project envelope · Current month 84% of cap
Within budget

GPU-hour pricing means you know the cost before you deploy. One audit trail tracks every request to the user who made it. You build. We govern.

Platform Preview

BESPIN WORKS OFFICIAL Governance Summary Hard Monthly Caps Strict budget ceilings per project to prevent cost overruns. Audit Trail Comprehensive logging of every request to user and project. Provider Neutrality Zero markup on provider keys; absolute vendor neutrality. Graduation Path Seamless transition from Sandbox to Production environments. GOVERNANCE CERTIFIED VP of Engineering / CTO Bespin Works Infrastructure

One endpoint, zero rebuilds

Start as an experiment, graduate to production — on the same endpoint. No rebuilding, no re-platforming, no new infrastructure.

Conceptual preview

# Case A: Using Bespin's managed compute
curl https://app.bespin.works/v1/chat/completions
  -H 'X-Bespin-Mode: operated'
  -H 'Content-Type: application/json'
  -d '{"model": "llama-3", "messages": [...]}'
# Case B: Using your own provider keys
curl https://app.bespin.works/v1/chat/completions
  -H 'X-Bespin-Mode: provider'
  -H 'Content-Type: application/json'
  -d '{"model": "gpt-4", "messages": [...]}'
# Result: Identical URL, different runtime execution.
1
Fit-Check
2
Governance Summary
3
Sandbox Access

Questions, answered

The short answers to the questions we hear most.

What is an AI Applications Platform?

An AI Applications Platform is the layer that runs, governs, and budgets enterprise AI applications. SaaS gives you the application runtime; GPU clouds give you the inference substrate; gateways route between models. Bespin Works combines all three, and prices by the project, not the token.

How is Bespin Works different from a token-based gateway?

Token-based gateways bill by the token, which makes invoices hard to predict. Bespin Works bills by GPU-hour and per project, with a hard monthly cap on every project, so you can model the cost before you deploy.

How does billing work?

Every project runs under a per-project plan with a hard monthly cap. Operated projects bill by GPU-hour; provider projects bill a flat per-project fee. Hit the cap and we throttle, not bill.

Do I need to own GPUs or build infrastructure?

No. In operated mode, Bespin Works runs open-weight models on managed compute. You write the application, we run the models. No GPU ownership, no infrastructure to provision.

Can I use my existing OpenAI, Anthropic, Azure, or Google keys?

Yes. In provider mode, you bring your own keys and Bespin Works routes, governs, and audits the traffic without touching, reselling, or marking up your provider spend.

What does "no vendor lock-in" mean?

You're not tied to a single model provider. Switch a project between operated compute and your own provider keys anytime, on the same endpoint, without re-platforming.

How does governance work, and what does it cover?

Every project gets a hard monthly cap, every request is attributed to a user, and every month you get a signable governance summary. Governance covers cost, access, and routing: one point of policy for all your enterprise AI.

How do I get started?

Talk to us. We'll discuss your AI application roadmap, understand your use case, and set up a fit-check. Email ceo@bespin.works.

Discuss your AI application roadmap

See how Bespin Works runs, governs, and budgets enterprise AI applications for your organization.

Privacy

What we do and don't collect when you visit this site.

We use Google Analytics and Microsoft Clarity to measure how visitors find and use this site. These services set cookies and similar technologies in your browser to report aggregate, anonymous traffic data and session recordings of how visitors move through the page. We use this to improve the site.

This site has no sign-up, no login, and no forms. The only way you contact us is by email, and we only use the email you send to reply to you.

Google and Microsoft are independent processors of the data their tools collect; their use is governed by their own privacy policies. You can decline or clear these cookies in your browser settings at any time, and the site continues to work without them.

Questions about privacy? Email ceo@bespin.works.