Qwen-Audio-3.0-TTS Online

Qwen Audio TTS — Generate Natural Speech with Qwen-Audio 3.0

Turn text into clear, expressive speech. With Qwen Audio TTS, enter a script, choose an available voice, and describe the delivery in ordinary language. Adjust pace, tone, emotion, or emphasis, then generate and preview the result in your browser.

The current tool uses the Qwen-Audio-3.0-TTS model family for voiceovers, explainers, localized content, product demos, learning materials, and conversational experiences—not music or sound-effect generation. Start with a short passage, compare the settings, and download the result that fits your project.

The live tool shows current models, voices, languages, input length, file options, and generation limits. Those values are the source of truth for this website.

Voice creator recording expressive narration in a modern studio

Speech-first workflow

Script, direction, voice, then a real generated result.

API contract connected

Qwen Audio TTS workspace

Flash and Plus configured
Model

Write the words your audience should hear.

0 charactersBilling is calculated per 100 characters.
Voice
Longan HuanSystem voice

28 curated voices for Flash. System voices appear first, followed by base voices selected from the official preview package.

System voice · Warm female voice · Social companion · Chinese / English

Describe tone, pace, emotion, and emphasis in ordinary language.

Advanced output controls

Flash and Longan Huan are an official model–voice combination. Current credits and backend limits are applied when you generate.

Output

Your generated speech appears here.

Waiting

No audio generated yet

Enter a script, choose a compatible voice, and generate a real audio result through the site API.

  • Successful results are added to your task history.
  • Voice cloning is not included in this first-release workflow.

Official base-voice examples

Hear Qwen Audio TTS in Action

Hear 24 official Qwen Audio TTS source previews before choosing a voice. As part of the Qwen3 TTS tool collection, this page uses selected audio from the base voice preview package linked in the Alibaba Cloud documentation. These are official fixed previews—not files generated by this website or results produced by the current controls.

Official preview

Longying Haikai

Upbeat, positive voice

Best fit
General customer service
Voice parameter
longyinghaikai
Models
Flash + Plus
Language
Chinese (Mandarin)
7.89s · Official base-voice packageSource
Official preview

Longying Haixuan

Capable, composed voice

Best fit
General customer service
Voice parameter
longyinghaixuan
Models
Flash + Plus
Language
Chinese (Mandarin)
9.81s · Official base-voice packageSource
Official preview

Longying Jingdong

Composed, professional voice

Best fit
Banking and credit service
Voice parameter
longyingjingdong
Models
Flash + Plus
Language
Chinese (Mandarin)
8.45s · Official base-voice packageSource
Official preview

Longying Haizhe

Rational, assertive voice

Best fit
Persuasive service
Voice parameter
longyinghaizhe
Models
Flash + Plus
Language
Chinese (Mandarin)
9.17s · Official base-voice packageSource
Official preview

Longying Jinhao

Approachable, reliable voice

Best fit
Community service
Voice parameter
longyingjinhao
Models
Flash + Plus
Language
Chinese (Mandarin)
8.45s · Official base-voice packageSource
Official preview

Longluan Xuanling

Gentle older-sister voice

Best fit
Audiobooks and knowledge sharing
Voice parameter
longluanxuanling
Models
Flash + Plus
Language
Chinese (Mandarin)
6.93s · Official base-voice packageSource
Official preview

Longyan Zhihe

Sweet, gentle voice

Best fit
Audiobooks and emotional companionship
Voice parameter
longyanzhihe
Models
Flash + Plus
Language
Chinese (Mandarin)
6.29s · Official base-voice packageSource
Official preview

Longhe Xiaoxuan

Refined, scholarly voice

Best fit
Audiobooks
Voice parameter
longhexiaoxuan
Models
Flash + Plus
Language
Chinese (Mandarin)
7.41s · Official base-voice packageSource
Official preview

Longsong Linwang

Warm, resonant voice

Best fit
Late-night radio and audiobooks
Voice parameter
longsonglinwang
Models
Flash + Plus
Language
Chinese (Mandarin)
6.53s · Official base-voice packageSource
Official preview

Longju Tongying

Sweet, playful voice

Best fit
Conversation and animation dubbing
Voice parameter
longjutongying
Models
Flash + Plus
Language
Chinese (Mandarin)
5.33s · Official base-voice packageSource
Official preview

Longrui Tinglin

Sweet, playful voice

Best fit
Conversation and animation dubbing
Voice parameter
longruitinglin
Models
Flash + Plus
Language
Chinese (Mandarin)
7.25s · Official base-voice packageSource
Official preview

Longhui Luling

Gentle, caring voice

Best fit
Conversation, companionship, and audiobooks
Voice parameter
longhuiluling
Models
Flash + Plus
Language
Chinese (Mandarin)
7.33s · Official base-voice packageSource
Official preview

Longshuo Jizhu

Standard broadcast voice

Best fit
News narration
Voice parameter
longshuojizhu
Models
Flash + Plus
Language
Chinese (Mandarin)
11.97s · Official base-voice packageSource
Official preview

Longyu Fengmo

Gentle, resilient voice

Best fit
Conversation and emotional companionship
Voice parameter
longyufengmo
Models
Flash + Plus
Language
Chinese (Mandarin)
5.81s · Official base-voice packageSource
Official preview

loongalicezhu

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongalicezhu
Models
Flash + Plus
Language
See source details
6.85s · Official base-voice packageSource
Official preview

loongliamli

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongliamli
Models
Flash + Plus
Language
See source details
7.17s · Official base-voice packageSource
Official preview

loongiriszhang

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongiriszhang
Models
Flash + Plus
Language
See source details
6.29s · Official base-voice packageSource
Official preview

loongowensun

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongowensun
Models
Flash + Plus
Language
See source details
5.17s · Official base-voice packageSource
Official preview

loongmasonlin

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongmasonlin
Models
Flash + Plus
Language
See source details
6.45s · Official base-voice packageSource
Official preview

loongsophiazhao

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongsophiazhao
Models
Flash + Plus
Language
See source details
4.85s · Official base-voice packageSource
Official preview

loongisaaczhao

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongisaaczhao
Models
Flash + Plus
Language
See source details
5.09s · Official base-voice packageSource
Official preview

loongwillowwu

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongwillowwu
Models
Flash + Plus
Language
See source details
5.97s · Official base-voice packageSource
Official preview

loongvioletzhang

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongvioletzhang
Models
Flash + Plus
Language
See source details
7.41s · Official base-voice packageSource
Official preview

loongvictorluo

Official base-voice preview

Best fit
Additional voice sample
Voice parameter
loongvictorluo
Models
Flash + Plus
Language
See source details
6.45s · Official base-voice packageSource

A Qwen Audio TTS base-voice suffix can be used with Flash or Plus by applying the corresponding model prefix. Use these previews to narrow the direction, then test your own script because wording and instructions still affect the final result.

Model overview

What Is Qwen Audio TTS?

Qwen Audio TTS is this page’s simple name for online speech generation powered by Qwen-Audio-3.0-TTS. Instead of only reading words neutrally, it accepts delivery directions such as calm and reassuring, energetic but controlled, or slower with emphasis on a product name.

The official technical overview presents Qwen-Audio-3.0-TTS as a production-oriented speech-synthesis family with Flash and Plus modes. Both support instruction control. Flash favors faster feedback; Plus favors professional-quality speech. The underlying Flash model supports voice cloning, but this website does not expose that feature in its first release. Neither mode supports voice design from text.

This is an independent third-party website. It is not an official Alibaba Cloud or Qwen service, and references to model names describe the underlying technology rather than an affiliation or endorsement.

Capability frame

Flash

Faster feedback for previews, iteration, and interactive speech.

Plus

Professional-quality speech for polished narration and final reads.

Instruction control

Direct tone, emotion, pacing, and emphasis in natural language.

Choose a mode

Qwen Audio TTS Flash vs Plus

Choose by project needs, not because one model name sounds universally better. Qwen Audio TTS Flash suits quick iteration and interactive speech. Plus is a better starting point when final output quality matters more than rapid experimentation.

CapabilityFlashPlus
Best fitFast previews, iterative work, interactive speechPolished narration and production-oriented speech
Natural-language delivery controlSupportedSupported
Voice cloningSupported by the underlying model, but not exposed in this website’s first releaseNot supported in current official documentation
Voice design from textNot supportedNot supported
Recommended workflowFind and refine a direction quicklyStart with a clear direction and optimize the final read

Qwen Audio TTS results vary with language, voice, script, and instructions. Test both modes with the same passage when available. The website does not promise a fixed response time because queues, input length, and network conditions also affect generation.

Four-step workflow

How to Use Qwen Audio TTS

01

Enter text written for listening

Paste a voiceover, dialogue, lesson, or product explanation. Punctuation shapes pauses and rhythm, so split crowded sentences before generating. The AI Text to Speech guide explains the basic text-first workflow.

02

Choose a model and voice

In the Qwen Audio TTS workspace, select Flash for faster iteration and interactive speech. Try Plus for a polished final read. Preview voices in your language instead of choosing only by name.

03

Describe the delivery

Write one concise direction for tone, pace, emotion, and emphasis. For example: “Warm and trustworthy, medium pace, pronounce every number clearly.” Add a verified inline tag only when one moment needs a local change.

04

Generate, preview, and download

Listen from beginning to end. Check names, numbers, abbreviations, pauses, and sentence endings. If something sounds wrong, revise the text or one instruction, then regenerate that section.

Qwen Audio TTS works best iteratively: make one change, listen again, and keep the version that improves the read.

Expression control

Control More Than the Words

The same sentence can sound reassuring, urgent, playful, restrained, or uncertain. Qwen Audio TTS lets you direct those differences without a large panel of audio-engineering controls.

Audio director refining narration tone and pacing from a script

Direct the voice in natural language

Write instructions as if briefing a voice actor. “Clear, patient, and slightly slower than normal” gives the model a coherent target. Avoid contradictory directions, and change one variable at a time.

Add local expression with inline tags

Official material describes 86 fine-grained inline tags across the model family, including emotion and non-verbal cues. A tag can shape one moment without changing an entire script. Only production-tested tags should appear in the live selector.

Shape pacing with punctuation and structure

Sentence length, commas, full stops, and paragraph breaks all influence how a script is delivered. Use shorter sentences for instructional content, add clear breaks before important ideas, and avoid excessive punctuation that creates unnatural pauses. Test one paragraph before applying the same structure to a longer script.

Work from short tests to longer speech

Official material describes single-pass synthesis up to about three minutes. That is a model capability, not a promise that every website plan exposes the full limit. Confirm the generator maximum and approve a short test first.

Multilingual reach

Languages and Chinese Dialects in Qwen Audio TTS

Official material describes 16 languages and 20 Chinese dialect regions across the Qwen-Audio-3.0-TTS family. That does not mean every preset voice speaks every language or that all options are enabled here.

The Qwen Audio TTS generator includes all six model-specific system voices and 24 curated base voices from the official preview package. Fourteen include the richer names and use cases published on the documentation page; the remaining previews keep their official suffix IDs. Search by name, suffix, style, or use case. Voice coverage can still differ by language and region, so treat the live voice and language controls as the availability list.

For multilingual work, have a fluent speaker review the script and test names, numbers, abbreviations, and borrowed words. Split frequent language switches into short sections. Qwen Audio TTS coverage is not a substitute for human review.

16

languages described across the model family

20

Chinese dialect regions described officially

Website availability can be narrower. The production menus—not these family-level figures—determine what users can generate.

Practical workflows

What Can You Create with Qwen Audio TTS?

Creators working on video narration, localization, voice assistants, and game characters

01

Video voiceovers and narration

Use Qwen Audio TTS to create speech for tutorials, product walkthroughs, documentary-style pieces, and social videos. Direct pacing and emphasis around the edit instead of reading every scene identically.

02

Localization and dubbing

Prepare spoken versions for different markets while keeping a consistent brand direction. Edit translations for natural speech, timing, and local context before generation.

03

Voice assistants and service prototypes

Generate prompts for demos, help centers, and conversational product flows. Flash suits quick feedback, but end-to-end responsiveness also depends on the surrounding application.

04

Games and original characters

Explore delivery styles for original characters, mission prompts, and narrative scenes. Use authorized characters and voices; do not imitate a real actor, public figure, or protected character voice.

Qwen Audio TTS is a speech tool. It should not be presented as a general generator for songs, background music, or standalone sound effects.

Script craft

How to Get Better Qwen Audio TTS Results

Warm and credible, medium pace. Pause after the first sentence and pronounce every product number clearly.

One coherent delivery direction

Write for the ear

Read the script aloud. Break long sentences, mark pauses with punctuation, and rewrite ambiguous abbreviations or symbols. Elegant written text can still sound crowded in one breath.

Give one clear direction

Use compatible qualities: “calm, credible, medium pace” is clearer than competing moods. If the result misses, change only pace, emotion, or emphasis and compare again.

Keep directions separate from spoken text

Put delivery guidance in the instruction field instead of mixing stage directions into the script. Otherwise, the model may speak words such as “pause” or “whisper.” Insert only inline tags that the live tool explicitly supports.

Preview difficult content first

Test names, product terms, dates, currency, acronyms, and mixed-language phrases first. Fix them before generating the complete narration.

Know the model family

Qwen Audio TTS vs Qwen3-TTS

The names are similar, but they should not be treated as interchangeable. Qwen Audio TTS on this page refers to the hosted Qwen-Audio-3.0-TTS family and its Flash/Plus service modes. Qwen3-TTS refers to a separate open-source model family with its own model sizes, deployment choices, and feature variants.

QuestionQwen-Audio-3.0-TTSQwen3-TTS
How this page uses itConnected online speech serviceSeparate tool and model family
Main choicesFlash and PlusModel size and capability variants
Voice cloningUnderlying Flash capability; not available on this page at launchDepends on the selected Qwen3-TTS variant
Voice designNot supported by Flash or PlusAvailable in specific Qwen3-TTS variants
Best decisionChoose by speed and final output qualityChoose by feature needs and deployment approach

Do not choose from the names alone. Use the current page for the Qwen-Audio-3.0-TTS online workflow. If you specifically need an existing cloning workflow, use the separate AI Voice Clone page. Review the official Qwen3-TTS repository for details about Qwen3-TTS variants or self-hosting.

Questions and limits

Qwen Audio TTS FAQ

Model-family facts and website availability are kept separate here. The live generator remains the source of truth for selectable options.

01. What is Qwen Audio TTS?

It is this website’s name for an online text-to-speech workflow powered by Qwen-Audio-3.0-TTS. It converts text into speech and supports natural-language delivery instructions through the options exposed in the generator.

02. Is Qwen Audio TTS the same as Qwen3-TTS?

No. Qwen-Audio-3.0-TTS and Qwen3-TTS are separate model families with different service modes and feature variants. This page identifies the current model explicitly so users can choose by capability rather than by a similar name.

03. Is this an official Qwen or Alibaba Cloud website?

No. This is an independent third-party tool and is not operated by or officially affiliated with Alibaba Cloud or Qwen. Model names are used to identify the technology connected to the experience.

04. What is the difference between Flash and Plus?

Flash emphasizes faster feedback, while Plus targets more professional-quality speech. The underlying Flash model supports voice cloning, but this page does not expose it at launch. Both support instruction control, and neither currently supports voice design from text.

05. Does Qwen Audio TTS support voice cloning?

The underlying Flash model supports voice cloning, but this website’s first release does not provide a reference-audio upload or cloning workflow. Plus does not support cloning in current official documentation.

06. Can it design a completely new voice from a description?

No. Current official documentation lists voice design as unsupported for both Qwen-Audio-3.0-TTS Flash and Plus. Choosing a preset voice and controlling its delivery is different from creating a new identity from text.

07. How many languages and dialects are supported?

Official model material describes 16 languages and 20 Chinese dialect regions across the family. Actual choices depend on mode, voice, and website integration, so check the live menus before planning a project.

08. How many Qwen Audio TTS voices are available here?

The selector includes all six system voices and 24 curated base voices from the official preview package. Fourteen have richer names and use cases published on the official voice-list page; the remaining previews use their official suffix IDs. A base suffix is combined with the selected Flash or Plus model to form the complete voice parameter.

09. Are all 86 inline tags available here?

Not necessarily. The number describes the official model family’s technical capability. Only tags shown in the tool and verified against the production API should be considered available on this website.

10. Can I generate three minutes of speech at once?

Official material describes up to about three minutes of single-pass model output. Website limits can be shorter because of plan, queue, or product constraints. Follow the current character and duration guidance in the generator.

11. Can it generate music or sound effects?

No. This page is designed for speech synthesis. It does not promise song generation, background music, standalone sound effects, or a general-purpose audio-production workflow.

12. Can I use generated speech commercially?

Commercial use depends on this website’s current Terms, the connected service terms, and your rights to every script, character, and reference voice. Qwen Audio TTS does not grant rights you do not already have.

13. Is it free to use?

Available credits and paid options can change. Check the generator and pricing page for current costs, limits, and any trial allowance before starting a long script.

Start with a short test

Generate Your First Qwen Audio Voice

Start with two or three sentences you know well. Choose a voice, describe one clear delivery direction, and listen before expanding the script. Qwen Audio TTS makes it easier to compare expressive speech options while keeping the workflow focused on the words your audience will actually hear.

Generate Speech