Alibaba’s audio-native model generates picture and synchronized soundtrack together — dialogue, ambience and effects in one pass.
HappyHorse 1.0 is Alibaba’s audio-native video model: instead of adding sound after the fact, it generates the soundtrack — speech, ambience, effects — together with the picture, keeping the two synchronized. That puts it in the same club as Google’s Veo for publish-ready clips.
On Nexvy, HappyHorse renders 5, 8 or 10-second clips at 720p or 1080p. It is a premium-priced model, so the typical workflow is to draft the scene on a cheaper model, then render the final with HappyHorse when synchronized audio matters.
Prices are in Nexvy credits — one balance shared by every model on the platform.
| Model | Resolution | Duration | Audio | Credits | Approx. price |
|---|---|---|---|---|---|
| HappyHorse 1.0 | up to 1080p | 5–10s | ✓ | 1620–4212 | $3.24–$8.42 |
Approximate USD price per video. Subscription plans lower the effective cost per credit.
01
Sign up in seconds — new accounts get free credits, no card required.
02
Open the video generator and choose HappyHorse 1.0 from the model selector.
03
Write a prompt or attach a start frame, then set duration and resolution.
04
Watch the clip render, iterate on the prompt and download the final video.
Characters whose speech is generated in sync with their motion.
Street scenes, nature and interiors with believable soundscapes.
Publishable video with sound, no audio post-production step.