Inference that costs 90% less.
Text to speech, chat, and batch inference on the best open models. Works with the OpenAI and ElevenLabs SDKs you already use. Change one line and keep everything else.
pip install openai−from openai import OpenAIclient = OpenAI(api_key=os.environ["MSERVE_API_KEY"],
)audio = client.audio.speech.create(model="kokoro-82m", voice="af_nicole",input="Your order is on its way.",)
Products
Speech, chat, and batch inference on open models.
Voice
Natural voices from the best open models. Clone a voice from a 10-second clip. Stream as you generate.
Real-time
Chat and completions on open models, through the OpenAI API you already use.
Batch
Send a file of requests and collect results within your window. Drop‑in for the OpenAI Batch API.
Speech that streams.
The first sentence plays while the rest renders.
Chat behind the OpenAI API.
Streaming replies from open models, same SDK.
Batch at our lowest price.
One file in, results within your window.
Your bill, after.
Voice cloning, per million characters. ElevenLabs Business plan, Flash models, elevenlabs.io/pricing, August 2026, against Chatterbox.
Voice cloning: $82.50 becomes $6.00. You save 93%.
Why teams switch.
Nothing to rewrite.
Keep your SDK, your code, and your prompts. Swap the base URL.
Prices you can see.
One public price per model. No seats, no quotes.
Credits that never expire.
Prepaid, billed per character, per token, or per minute. Purchased credits never expire. No monthly plan.
Models you can build a business on.
Apache 2.0 and MIT licensed. Commercial use included.
Models and pricing
| Model | Product | Best for | License | Price |
|---|---|---|---|---|
Kokoro 82M | Voice | Preset voices | Apache-2.0 | $3.00/ 1M chars |
Chatterbox | Voice | Voice cloning | MIT | $6.00/ 1M chars |
Chatterbox Voice Conversion | Voice | Voice conversion | MIT | $0.06/ minute |
Whisper Large v3 Turbo | Speech to text | Transcription | MIT | $0.006/ minute |
Qwen3.8 27B | Real-time | Chat | Apache-2.0 | Input$0.30/ 1M tokens Output$2.50/ 1M tokens |
Qwen3.8 27B (Batch) | Batch | Batch jobs | Apache-2.0 | 24 h$0.22/ 1M tokens Flex$0.15/ 1M tokens |
Qwen3 Embedding 0.6B | Embeddings | Search, retrieval | Apache-2.0 | $0.01/ 1M tokens |
FAQ
Does it work with the OpenAI SDK?
Yes. Chat, completions, responses, embeddings, speech, transcription, and batch work with the official openai libraries. Set base_url and your key.
Does it work with the ElevenLabs SDK?
Yes. Text to speech, streaming, voice cloning, and speech to text work with the official elevenlabs library. ElevenLabs premade voice ids work unchanged.
How is it this cheap?
We run open models on efficient hardware and pass the savings on. Every price is public.
Is there a free tier?
Add a card and get $5 of credit, valid for 30 days. No charge until you top up.
When can I get access?
We onboard from the waitlist. Join and we email you when your key is ready.
Can I use the voices commercially?
Yes. Every model we serve is Apache 2.0 or MIT licensed.
Start building for less.
One email when your key is ready. Nothing else.