TTS and Voice Announcements: Methods, Engines, Language, Volume, Delay, and Quality
TTS and Voice Announcements: Methods, Engines, Language, Volume, Delay, and Quality
We're back to the script.home_announcement skeleton from Lesson 1. Today we add a real TTS (text-to-speech) engine to it: converting text into speech played through a chosen speaker.
Budget around 80 minutes. Example entities: tts.piper or tts.google_translate_en, media_player.living_room_speaker.
Choosing a TTS Engine: Local Versus Cloud
Home Assistant supports several TTS engines. Piper is a local neural engine that runs offline, with no text sent to any cloud: good voice quality, zero network latency, full privacy. Cloud alternatives (integrations built on Google Translate TTS or commercial engines, for example) can sound a bit more natural, but need the internet and send your announcement's content outside the home network, which matters for critical announcements: if the internet drops right at the moment of a flood, a cloud engine won't speak the warning.
Recommendation: Piper for critical announcements, any engine for home ones
Configure Piper (local, offline) as the engine for script.critical_announcement, so internet dependency never blocks a flood or smoke warning. For home announcements (less urgent), you can use any engine, cloud-based included, if you prefer its voice.
The Full home_announcement Script
script:
home_announcement:
alias: "Home announcement"
icon: mdi:bullhorn-outline
mode: queued
max: 5
fields:
message:
description: "The message to speak"
example: "Laundry is done"
sequence:
- condition: state
entity_id: input_boolean.house_silence
state: "off"
- condition: state
entity_id: input_boolean.multimedia_automation_enabled
state: "on"
- variables:
zone_speakers: >
{% if is_state('input_select.announcement_zone', 'Whole house') %}
media_player.living_room_speaker,media_player.kitchen_speaker,media_player.bedroom_speaker
{% elif is_state('input_select.announcement_zone', 'Downstairs') %}
media_player.living_room_speaker,media_player.kitchen_speaker
{% else %}
media_player.living_room_speaker
{% endif %}
- action: media_player.volume_set
target:
entity_id: "{{ zone_speakers.split(',') }}"
data:
volume_level: "{{ states('input_number.announcement_volume_home') | float / 100 }}"
- action: tts.speak
target:
entity_id: tts.piper
data:
media_player_entity_id: "{{ zone_speakers.split(',') }}"
message: "{{ message }}"
language: "en"
The script reads the announcement zone from the Lesson 1 helper, sets the volume to the value from input_number.announcement_volume_home (dividing by 100, since you'll remember from Lesson 2 that volume_level is a 0.0-1.0 range), and calls tts.speak on the chosen speakers. language: "en" matters: without explicitly specifying the language, some engines may try to guess it from the text and mispronounce names or technical terms.
The critical_announcement Script: Overriding Silence
script:
critical_announcement:
alias: "Critical announcement"
icon: mdi:alert-octagon
mode: parallel
fields:
message:
description: "The critical message to speak"
example: "Flood detected in the bathroom"
sequence:
- action: media_player.volume_set
target:
entity_id:
- media_player.living_room_speaker
- media_player.kitchen_speaker
- media_player.bedroom_speaker
data:
volume_level: "{{ states('input_number.announcement_volume_critical') | float / 100 }}"
- action: tts.speak
target:
entity_id: tts.piper
data:
media_player_entity_id:
- media_player.living_room_speaker
- media_player.kitchen_speaker
- media_player.bedroom_speaker
message: "Attention. {{ message }}. Repeating. {{ message }}."
language: "en"
This script deliberately doesn't check input_boolean.house_silence: a critical announcement has every right to override silence, because the cost of missing a flood alert is higher than the cost of waking someone up. Repeating the message ("Repeating...") compensates for the fact that the first few words can be missed while someone is asleep or focused on something else.
Delay: Why the First Word Sometimes Gets Lost
A speaker needs a moment to "wake up" its audio
Some speakers (especially power-saving ones) have a short delay between receiving a command and actually starting to play sound: the first half-second of an announcement can get "clipped." If you notice this on your own hardware, add a short signal chime before the actual TTS announcement, so the speaker has time to fully spin up before the important content plays.
TTS Engines: Piper Versus Cloud, and Dedicated Hardware
| Engine | Type | Characteristics |
|---|---|---|
| Piper | local | designed for weaker hardware, generates speech faster than real time even on a Raspberry Pi 4; natural-sounding voices, though not quite at commercial-cloud level |
| Google Cloud TTS / Amazon Polly / Azure | cloud | very good voice quality, but needs an account, an API key, and usually per-character billing |
| Home Assistant Voice Preview Edition | dedicated device | Nabu Casa hardware with microphones and an XMOS audio processor, can run fully locally or with Nabu Casa's privacy-focused cloud |
Since October 2025, Piper has been developed as the OHF-Voice/piper1-gpl project (the original rhasspy/piper repository was archived), so if you're looking for documentation, check the current source. For most home announcements, status updates, alerts, notifications, Piper is entirely sufficient: the quality gap versus cloud engines matters mainly for longer, natural conversation, not short alerts.
If you're planning to expand into a full voice assistant (not just outgoing announcements, but incoming voice commands too), it's worth looking early at Wyoming satellite-class devices: the Home Assistant Voice Preview Edition, or your own build on an ESP32-S3 with ESPHome's voice components, which now support even two wake words at once on a single device.
Multiple Languages and Accents in Voice Announcements
If more than one language is spoken in your house, or your integrations generate announcements containing foreign-language names (track titles, device brand names, for example), it's worth checking whether your chosen TTS engine pronounces mixed text correctly. Piper, the local TTS engine covered in this lesson, needs a separate voice model for each language, so an announcement mixing two languages in one sentence can sound unnatural if both parts get synthesized by the same language model.
Voice Caching: Saving Resources on Repeated Announcements
Generating speech from text (especially through cloud engines) takes time and, for paid services, generates a cost on every call. Home Assistant caches (stores) generated audio files for identical text by default, so a recurring announcement, like "Gate open," gets generated only once, and later playbacks reuse the saved file without going back through the TTS engine. This noticeably cuts reaction time for frequently repeated security-automation announcements.
Offline TTS on Weaker Hardware: Optimization Strategies
A local TTS engine like Piper needs real compute, which on weaker hardware (an older Raspberry Pi simultaneously running many other integrations from this course, say) can introduce a noticeable delay before an announcement starts. The fix is either picking a smaller, faster Piper voice model (at the cost of a slightly less natural sound), or moving TTS processing to dedicated, more powerful hardware, like the Home Assistant Voice Preview Edition mentioned in this lesson, with its own audio processor.
Neural Voices Versus Classic Speech Synthesis Engines
Older TTS engines (like espeak) generate a clearly artificial, robotic voice, while newer neural models (Piper included) sound noticeably more natural, at the cost of higher compute requirements. For security announcements, where clarity matters more than natural sound, the difference matters less, but for everyday informational announcements, heard many times a day, a more natural neural voice noticeably improves how comfortable the system is to live with.
Personalizing the TTS Voice to Fit Your Home's Identity
Piper and other TTS engines offer a choice among several different voices for the same language, differing in tone and speaking pace. It's worth spending a moment picking a voice that actually fits your home's character and your household's preferences, instead of leaving the default setting: a warmer, slower voice for evening announcements, a more energetic one for morning reminders, if the engine allows it.
Testing Announcement Clarity in a Noisy Environment
An announcement that sounds clear in a quiet room can become unintelligible in a kitchen with a range hood running, or a room with a vacuum going. It's worth testing key security announcements under realistically noisy conditions, not just in silence, and adjusting the critical_announcement script's base volume if needed so it cuts through typical household noise, not just in ideal test conditions.
The Cost of Cloud TTS Engines at High Announcement Volume
Cloud-based TTS engines (like some Google or Amazon services) are usually billed per character converted to speech, which, under very heavy use of home announcements, can add up to a noticeable monthly cost. The local Piper engine, while it needs its own compute, eliminates that cost entirely, which makes it an attractive option not just for privacy, but for long-term economy with the kind of extensive announcement system this module builds.
How to Test This Lesson
1. Call script.home_announcement with a sample message and check it picks the right speakers for the chosen zone.
2. Turn on input_boolean.house_silence and confirm the home announcement doesn't play, but the critical one still does.
3. Check names, addresses, or technical terms in a spoken message: does the engine pronounce them correctly.
Common Mistakes
A critical announcement on a cloud engine: no internet blocks the warning at the worst possible moment.
No explicit language: "en": the engine guesses the language and mispronounces names.
Silence blocking a critical announcement: the input_boolean.house_silence condition accidentally copied into the wrong script.
Practical Task
☐ Install and configure Piper as your local TTS engine.
☐ Build both scripts: home_announcement and critical_announcement.
☐ Test both in silence mode to confirm the different behavior.
Key Takeaways
Offline Piper for critical announcements: independence from the internet.
The critical announcement deliberately bypasses silence mode, the home announcement respects it.
language: "en" always set explicitly, to avoid mispronunciation.
What's Next
Next lesson: Announcements Without Ruining the Music: Snapshot, Pause/Resume, Music Assistant, and Priority. What happens when an announcement hits a speaker that's playing music.