OpenAI’s audio models act on stage directions: tone, pacing, accents and character all steerable from the prompt.
GPT Audio is OpenAI’s family of speech-generation models. Their strength is direction-following: describe how the line should be delivered — "whisper it conspiratorially", "upbeat radio host", "slow down on the last sentence" — and the model acts it out, rather than just reading the text.
Nexvy carries three tiers: GPT Audio (best quality, 160 credits), GPT Audio Mini (fast and cheap at 40 credits — ideal for drafts and volume) and GPT-4o Audio (the legacy conversational model). They complement ElevenLabs: GPT Audio for acted delivery, ElevenLabs for polished narration voices.
Prices are in Nexvy credits — one balance shared by every model on the platform.
| Model | Resolution | Credits | Approx. price |
|---|---|---|---|
| GPT Audio | — | 160 | $0.32 |
| GPT Audio Mini | — | 40 | $0.08 |
| GPT-4o Audio | — | 160 | $0.32 |
Approximate USD price per generation. Subscription plans lower the effective cost per credit.
01
Sign up in seconds — new accounts get free credits, no card required.
02
Open the audio generator and choose GPT Audio as your model.
03
Paste a script and pick a voice, or describe the sound you need.
04
Listen to the result, tweak delivery or wording, and export the audio file.
Acted lines with emotion and personality for games, ads and animation.
Mini reads whole scripts for timing checks at 40 credits a pass.
Fine control over pacing and emphasis via natural-language notes.