Qwen-Audio-3.0-TTS Online
Qwen Audio TTS — Generate Natural Speech with Qwen-Audio 3.0
Turn text into clear, expressive speech. With Qwen Audio TTS, enter a script, choose an available voice, and describe the delivery in ordinary language. Adjust pace, tone, emotion, or emphasis, then generate and preview the result in your browser.
The current tool uses the Qwen-Audio-3.0-TTS model family for voiceovers, explainers, localized content, product demos, learning materials, and conversational experiences—not music or sound-effect generation. Start with a short passage, compare the settings, and download the result that fits your project.
The live tool shows current models, voices, languages, input length, file options, and generation limits. Those values are the source of truth for this website.

Speech-first workflow
Script, direction, voice, then a real generated result.
API contract connected
Qwen Audio TTS workspace
Output
Your generated speech appears here.
No audio generated yet
Enter a script, choose a compatible voice, and generate a real audio result through the site API.
- Successful results are added to your task history.
- Voice cloning is not included in this first-release workflow.
Official base-voice examples
Hear Qwen Audio TTS in Action
Hear 24 official Qwen Audio TTS source previews before choosing a voice. As part of the Qwen3 TTS tool collection, this page uses selected audio from the base voice preview package linked in the Alibaba Cloud documentation. These are official fixed previews—not files generated by this website or results produced by the current controls.
Longying Haikai
Upbeat, positive voice
- Best fit
- General customer service
- Voice parameter
- longyinghaikai
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longying Haixuan
Capable, composed voice
- Best fit
- General customer service
- Voice parameter
- longyinghaixuan
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longying Jingdong
Composed, professional voice
- Best fit
- Banking and credit service
- Voice parameter
- longyingjingdong
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longying Haizhe
Rational, assertive voice
- Best fit
- Persuasive service
- Voice parameter
- longyinghaizhe
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longying Jinhao
Approachable, reliable voice
- Best fit
- Community service
- Voice parameter
- longyingjinhao
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longluan Xuanling
Gentle older-sister voice
- Best fit
- Audiobooks and knowledge sharing
- Voice parameter
- longluanxuanling
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longyan Zhihe
Sweet, gentle voice
- Best fit
- Audiobooks and emotional companionship
- Voice parameter
- longyanzhihe
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longhe Xiaoxuan
Refined, scholarly voice
- Best fit
- Audiobooks
- Voice parameter
- longhexiaoxuan
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longsong Linwang
Warm, resonant voice
- Best fit
- Late-night radio and audiobooks
- Voice parameter
- longsonglinwang
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longju Tongying
Sweet, playful voice
- Best fit
- Conversation and animation dubbing
- Voice parameter
- longjutongying
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longrui Tinglin
Sweet, playful voice
- Best fit
- Conversation and animation dubbing
- Voice parameter
- longruitinglin
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longhui Luling
Gentle, caring voice
- Best fit
- Conversation, companionship, and audiobooks
- Voice parameter
- longhuiluling
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longshuo Jizhu
Standard broadcast voice
- Best fit
- News narration
- Voice parameter
- longshuojizhu
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
Longyu Fengmo
Gentle, resilient voice
- Best fit
- Conversation and emotional companionship
- Voice parameter
- longyufengmo
- Models
- Flash + Plus
- Language
- Chinese (Mandarin)
loongalicezhu
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongalicezhu
- Models
- Flash + Plus
- Language
- See source details
loongliamli
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongliamli
- Models
- Flash + Plus
- Language
- See source details
loongiriszhang
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongiriszhang
- Models
- Flash + Plus
- Language
- See source details
loongowensun
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongowensun
- Models
- Flash + Plus
- Language
- See source details
loongmasonlin
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongmasonlin
- Models
- Flash + Plus
- Language
- See source details
loongsophiazhao
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongsophiazhao
- Models
- Flash + Plus
- Language
- See source details
loongisaaczhao
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongisaaczhao
- Models
- Flash + Plus
- Language
- See source details
loongwillowwu
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongwillowwu
- Models
- Flash + Plus
- Language
- See source details
loongvioletzhang
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongvioletzhang
- Models
- Flash + Plus
- Language
- See source details
loongvictorluo
Official base-voice preview
- Best fit
- Additional voice sample
- Voice parameter
- loongvictorluo
- Models
- Flash + Plus
- Language
- See source details
A Qwen Audio TTS base-voice suffix can be used with Flash or Plus by applying the corresponding model prefix. Use these previews to narrow the direction, then test your own script because wording and instructions still affect the final result.
Model overview
What Is Qwen Audio TTS?
Qwen Audio TTS is this page’s simple name for online speech generation powered by Qwen-Audio-3.0-TTS. Instead of only reading words neutrally, it accepts delivery directions such as calm and reassuring, energetic but controlled, or slower with emphasis on a product name.
The official technical overview presents Qwen-Audio-3.0-TTS as a production-oriented speech-synthesis family with Flash and Plus modes. Both support instruction control. Flash favors faster feedback; Plus favors professional-quality speech. The underlying Flash model supports voice cloning, but this website does not expose that feature in its first release. Neither mode supports voice design from text.
This is an independent third-party website. It is not an official Alibaba Cloud or Qwen service, and references to model names describe the underlying technology rather than an affiliation or endorsement.
Capability frame
Flash
Faster feedback for previews, iteration, and interactive speech.
Plus
Professional-quality speech for polished narration and final reads.
Instruction control
Direct tone, emotion, pacing, and emphasis in natural language.
Choose a mode
Qwen Audio TTS Flash vs Plus
Choose by project needs, not because one model name sounds universally better. Qwen Audio TTS Flash suits quick iteration and interactive speech. Plus is a better starting point when final output quality matters more than rapid experimentation.
| Capability | Flash | Plus |
|---|---|---|
| Best fit | Fast previews, iterative work, interactive speech | Polished narration and production-oriented speech |
| Natural-language delivery control | Supported | Supported |
| Voice cloning | Supported by the underlying model, but not exposed in this website’s first release | Not supported in current official documentation |
| Voice design from text | Not supported | Not supported |
| Recommended workflow | Find and refine a direction quickly | Start with a clear direction and optimize the final read |
Qwen Audio TTS results vary with language, voice, script, and instructions. Test both modes with the same passage when available. The website does not promise a fixed response time because queues, input length, and network conditions also affect generation.
Four-step workflow
How to Use Qwen Audio TTS
Enter text written for listening
Paste a voiceover, dialogue, lesson, or product explanation. Punctuation shapes pauses and rhythm, so split crowded sentences before generating. The AI Text to Speech guide explains the basic text-first workflow.
Choose a model and voice
In the Qwen Audio TTS workspace, select Flash for faster iteration and interactive speech. Try Plus for a polished final read. Preview voices in your language instead of choosing only by name.
Describe the delivery
Write one concise direction for tone, pace, emotion, and emphasis. For example: “Warm and trustworthy, medium pace, pronounce every number clearly.” Add a verified inline tag only when one moment needs a local change.
Generate, preview, and download
Listen from beginning to end. Check names, numbers, abbreviations, pauses, and sentence endings. If something sounds wrong, revise the text or one instruction, then regenerate that section.
Qwen Audio TTS works best iteratively: make one change, listen again, and keep the version that improves the read.
Expression control
Control More Than the Words
The same sentence can sound reassuring, urgent, playful, restrained, or uncertain. Qwen Audio TTS lets you direct those differences without a large panel of audio-engineering controls.

Direct the voice in natural language
Write instructions as if briefing a voice actor. “Clear, patient, and slightly slower than normal” gives the model a coherent target. Avoid contradictory directions, and change one variable at a time.
Add local expression with inline tags
Official material describes 86 fine-grained inline tags across the model family, including emotion and non-verbal cues. A tag can shape one moment without changing an entire script. Only production-tested tags should appear in the live selector.
Shape pacing with punctuation and structure
Sentence length, commas, full stops, and paragraph breaks all influence how a script is delivered. Use shorter sentences for instructional content, add clear breaks before important ideas, and avoid excessive punctuation that creates unnatural pauses. Test one paragraph before applying the same structure to a longer script.
Work from short tests to longer speech
Official material describes single-pass synthesis up to about three minutes. That is a model capability, not a promise that every website plan exposes the full limit. Confirm the generator maximum and approve a short test first.
Multilingual reach
Languages and Chinese Dialects in Qwen Audio TTS
Official material describes 16 languages and 20 Chinese dialect regions across the Qwen-Audio-3.0-TTS family. That does not mean every preset voice speaks every language or that all options are enabled here.
The Qwen Audio TTS generator includes all six model-specific system voices and 24 curated base voices from the official preview package. Fourteen include the richer names and use cases published on the documentation page; the remaining previews keep their official suffix IDs. Search by name, suffix, style, or use case. Voice coverage can still differ by language and region, so treat the live voice and language controls as the availability list.
For multilingual work, have a fluent speaker review the script and test names, numbers, abbreviations, and borrowed words. Split frequent language switches into short sections. Qwen Audio TTS coverage is not a substitute for human review.
16
languages described across the model family
20
Chinese dialect regions described officially
Practical workflows
What Can You Create with Qwen Audio TTS?

01
Video voiceovers and narration
Use Qwen Audio TTS to create speech for tutorials, product walkthroughs, documentary-style pieces, and social videos. Direct pacing and emphasis around the edit instead of reading every scene identically.
02
Localization and dubbing
Prepare spoken versions for different markets while keeping a consistent brand direction. Edit translations for natural speech, timing, and local context before generation.
03
Voice assistants and service prototypes
Generate prompts for demos, help centers, and conversational product flows. Flash suits quick feedback, but end-to-end responsiveness also depends on the surrounding application.
04
Games and original characters
Explore delivery styles for original characters, mission prompts, and narrative scenes. Use authorized characters and voices; do not imitate a real actor, public figure, or protected character voice.
Qwen Audio TTS is a speech tool. It should not be presented as a general generator for songs, background music, or standalone sound effects.
Script craft
How to Get Better Qwen Audio TTS Results
Warm and credible, medium pace. Pause after the first sentence and pronounce every product number clearly.
One coherent delivery direction
Write for the ear
Read the script aloud. Break long sentences, mark pauses with punctuation, and rewrite ambiguous abbreviations or symbols. Elegant written text can still sound crowded in one breath.
Give one clear direction
Use compatible qualities: “calm, credible, medium pace” is clearer than competing moods. If the result misses, change only pace, emotion, or emphasis and compare again.
Keep directions separate from spoken text
Put delivery guidance in the instruction field instead of mixing stage directions into the script. Otherwise, the model may speak words such as “pause” or “whisper.” Insert only inline tags that the live tool explicitly supports.
Preview difficult content first
Test names, product terms, dates, currency, acronyms, and mixed-language phrases first. Fix them before generating the complete narration.
Know the model family
Qwen Audio TTS vs Qwen3-TTS
The names are similar, but they should not be treated as interchangeable. Qwen Audio TTS on this page refers to the hosted Qwen-Audio-3.0-TTS family and its Flash/Plus service modes. Qwen3-TTS refers to a separate open-source model family with its own model sizes, deployment choices, and feature variants.
| Question | Qwen-Audio-3.0-TTS | Qwen3-TTS |
|---|---|---|
| How this page uses it | Connected online speech service | Separate tool and model family |
| Main choices | Flash and Plus | Model size and capability variants |
| Voice cloning | Underlying Flash capability; not available on this page at launch | Depends on the selected Qwen3-TTS variant |
| Voice design | Not supported by Flash or Plus | Available in specific Qwen3-TTS variants |
| Best decision | Choose by speed and final output quality | Choose by feature needs and deployment approach |
Do not choose from the names alone. Use the current page for the Qwen-Audio-3.0-TTS online workflow. If you specifically need an existing cloning workflow, use the separate AI Voice Clone page. Review the official Qwen3-TTS repository for details about Qwen3-TTS variants or self-hosting.
Questions and limits
Qwen Audio TTS FAQ
Model-family facts and website availability are kept separate here. The live generator remains the source of truth for selectable options.
01. What is Qwen Audio TTS?
It is this website’s name for an online text-to-speech workflow powered by Qwen-Audio-3.0-TTS. It converts text into speech and supports natural-language delivery instructions through the options exposed in the generator.
02. Is Qwen Audio TTS the same as Qwen3-TTS?
No. Qwen-Audio-3.0-TTS and Qwen3-TTS are separate model families with different service modes and feature variants. This page identifies the current model explicitly so users can choose by capability rather than by a similar name.
03. Is this an official Qwen or Alibaba Cloud website?
No. This is an independent third-party tool and is not operated by or officially affiliated with Alibaba Cloud or Qwen. Model names are used to identify the technology connected to the experience.
04. What is the difference between Flash and Plus?
Flash emphasizes faster feedback, while Plus targets more professional-quality speech. The underlying Flash model supports voice cloning, but this page does not expose it at launch. Both support instruction control, and neither currently supports voice design from text.
05. Does Qwen Audio TTS support voice cloning?
The underlying Flash model supports voice cloning, but this website’s first release does not provide a reference-audio upload or cloning workflow. Plus does not support cloning in current official documentation.
06. Can it design a completely new voice from a description?
No. Current official documentation lists voice design as unsupported for both Qwen-Audio-3.0-TTS Flash and Plus. Choosing a preset voice and controlling its delivery is different from creating a new identity from text.
07. How many languages and dialects are supported?
Official model material describes 16 languages and 20 Chinese dialect regions across the family. Actual choices depend on mode, voice, and website integration, so check the live menus before planning a project.
08. How many Qwen Audio TTS voices are available here?
The selector includes all six system voices and 24 curated base voices from the official preview package. Fourteen have richer names and use cases published on the official voice-list page; the remaining previews use their official suffix IDs. A base suffix is combined with the selected Flash or Plus model to form the complete voice parameter.
09. Are all 86 inline tags available here?
Not necessarily. The number describes the official model family’s technical capability. Only tags shown in the tool and verified against the production API should be considered available on this website.
10. Can I generate three minutes of speech at once?
Official material describes up to about three minutes of single-pass model output. Website limits can be shorter because of plan, queue, or product constraints. Follow the current character and duration guidance in the generator.
11. Can it generate music or sound effects?
No. This page is designed for speech synthesis. It does not promise song generation, background music, standalone sound effects, or a general-purpose audio-production workflow.
12. Can I use generated speech commercially?
Commercial use depends on this website’s current Terms, the connected service terms, and your rights to every script, character, and reference voice. Qwen Audio TTS does not grant rights you do not already have.
13. Is it free to use?
Available credits and paid options can change. Check the generator and pricing page for current costs, limits, and any trial allowance before starting a long script.
Start with a short test
Generate Your First Qwen Audio Voice
Start with two or three sentences you know well. Choose a voice, describe one clear delivery direction, and listen before expanding the script. Qwen Audio TTS makes it easier to compare expressive speech options while keeping the workflow focused on the words your audience will actually hear.