How to Record a School Robocall
District recordings sound bad for two reasons that have nothing to do with the script. The first is that a phone line throws away most of the audio you carefully recorded. The second is that the bilingual message loses its Spanish-speaking listeners in the first ten seconds of English. Both are fixable in an afternoon.
| Setting | Do this | Because |
|---|---|---|
| Microphone | Any USB cardioid mic, or a wired headset. Not the laptop's built-in mic | A laptop mic captures the room, and the room is what ruins it |
| Distance | A hand-span from the mouth, slightly off-axis | Closer gives boom and plosives; further gives room |
| Room | Small, carpeted, soft furnishings, curtains drawn. A stationery cupboard beats a boardroom | Reverberation survives the phone line and destroys intelligibility |
| Background | HVAC off, phones off, corridor door shut, sign on the door | Steady noise raises the floor and eats consonants |
| Sample rate | Record at 44.1 kHz; it will be downsampled | Record clean, let the system reduce it |
| Bandwidth reality | The listener hears roughly 300 Hz to 3.4 kHz, mono | Everything below and above is discarded — the warmth you liked is gone |
| Level | Peaks around −6 dB, consistent, no clipping | Quiet recordings get amplified along with the noise |
| Compression | Light, if available | A phone line has little dynamic range; even levels survive it |
| Check on | An actual mobile phone, on speaker, in a car park | Laptop speakers flatter a recording; a mobile in wind does not |
The bandwidth row is the one that changes behaviour. A voice call is narrowband and mono: everything under about 300 Hz and over about 3.4 kHz is thrown away. This is why a recording that sounded rich and authoritative in the conference room sounds thin and mumbled on a parent's phone, and why clear consonants matter far more than a pleasant voice.
Pacing, and the stopwatch method
Read the script aloud, standing, with a stopwatch, twice. Note the time. That is your rate, and it is probably faster than you think under recording conditions and slower than you think when you are nervous.
Word budgets for 60 seconds: 140 words at 140 wpm, 150 at 150, 160 at 160. Use 140 for anything emergency-related and 150 for routine operations. Anything above 160 wpm is difficult to follow on a narrowband line and materially harder for a listener whose first language is not English.
Four pacing habits that survive the phone line:
- Pause a full beat between the headline and the detail. On a phone this reads as structure. Without it the message is a wall.
- Slow down for numbers. Times, dates, addresses and phone numbers should be read at roughly half your normal rate, with a pause before and after.
- Say phone numbers in groups with pauses, and say them twice if the message asks people to call.
- Do not smile through an emergency message. The convention that a smile improves a recording comes from customer service and is wrong for a lockdown.
Human voice or text-to-speech
The honest answer is that both are right, for different messages.
Use a human voice for: any emergency or incident; any message where the district's credibility is the point; the all-clear; anything about a death; the first message of the school year; and any message a family may hear twice and remember.
Use text-to-speech for: overnight and pre-dawn sends where nobody is available to record; high-frequency operational messages such as daily absence notifications; languages where you have no fluent reader on staff, because a competent synthetic read in Vietnamese is better than a phonetic attempt by an English speaker; and messages assembled from data where a human read would be error-prone.
Modern synthetic voices are good enough that most listeners will not identify them on a phone line, which is exactly why the distinction should be a policy rather than a habit. Districts drift toward text-to-speech for everything because it is faster, and then find that the lockdown message sounded like the book fair message.
One rule worth writing into the policy: the superintendent records the emergency messages. Not the communications director, not a synthetic voice. When something serious has happened, the voice families hear should be the one accountable for it, and the recording should be made freshly rather than pulled from a library — a pre-recorded emergency message will eventually be sent for the wrong incident.
Names, and when to leave them out
Names are where recordings do the most avoidable damage. Three cases:
- Staff names. Ask the person how they say their own name, record the answer, keep a pronunciation list in the communications folder. Getting a colleague's name wrong on a message to twelve thousand households is a small humiliation delivered at scale.
- School and place names. Local pronunciations of street and school names are frequently not what spelling suggests, and a superintendent new to the district will get one wrong. Have someone local listen before it goes out.
- Family and student names. A broadcast should not contain one. If a message concerns an individual child, it is a call to that family, not a broadcast — and the merge-field version of this is worse, because a recorded message that splices in a synthesised student name will mangle a proportion of them and each mangling lands with the family least likely to forgive it.
Where a name would be mispronounced and there is no time to check, use the role: "the school office", "your child's teacher", "the attendance office". A correct role beats a mangled name every time.
Bilingual recordings
The single most common defect in district voice broadcasts: sixty seconds of English followed by sixty seconds of Spanish, and the Spanish-speaking listener hangs up at fifteen seconds because they have no reason to believe anything is coming for them.
Three workable patterns:
- Separate sends by language of record. Best when your language data is good. Each family hears one message in their language, the call is half as long, and the cost per family is the same. This requires reliable language-of-record data, which many districts do not have.
- Bridged bilingual, one call. Open with a five-second bilingual bridge before the English begins:
"This message is in English and Spanish. Este mensaje está en inglés y español. El español comienza en un minuto."That one sentence is the difference between a listener waiting and a listener hanging up. Then English, then a two-second gap, then Spanish. - Spanish first. In a district where Spanish-preference households are the majority, put Spanish first and bridge in English. Districts rarely consider this and it is often correct.
Use two voices, one per language, each a fluent speaker. A single reader working through a language they do not speak is worse than a synthetic voice, and families can tell.
What the platform contributes, and what it does not. Kastr's translation preview renders a draft through DeepL in up to five languages before you send, which is where you catch a mistranslated bus instruction — that is a text-side check, and it does not review your recording. Translations are cached on a hash of source text and language pair, so a district never pays twice for the same closure sentence, and the cache holds no index tying a string to a family. What we do not have: a translation glossary, a "see original" footer on translated messages, and any AI drafting assistance. If a vendor tells you their platform maintains your district's terminology glossary, ask to see it configured.
Questions people actually ask
Should a school robocall use a real voice or text to speech?
Human for emergencies, incidents, all-clears and anything where the district's credibility is the message — and the superintendent should be the voice on those. Text-to-speech for pre-dawn sends, high-frequency operational messages, and languages where you have no fluent reader on staff. Make it a written policy, because the default drift is toward synthetic for everything.
How fast should you speak on a recorded school phone message?
140 words per minute for anything emergency-related and around 150 for routine operations, which puts a 60-second message at 135 to 150 words. Slow to roughly half rate for times, dates, addresses and phone numbers, and pause a full beat between the headline and the detail so the message has audible structure.
What audio quality do phone broadcasts actually deliver?
Narrowband mono, roughly 300 Hz to 3.4 kHz. Everything below and above is discarded, which is why a recording that sounds warm on laptop speakers sounds thin on a phone. Record in a small soft room with a cardioid mic at a hand-span, keep levels even, and check the result on an actual mobile on speaker before you send it.
Who should record the district's emergency messages?
The superintendent, freshly, for each incident. Families should hear the voice of the person accountable, and a pre-recorded emergency library will eventually be sent for the wrong incident. Keep a named deputy in the policy for the case where the superintendent is inside the incident or unreachable.
One price. Every feature. Locked for three years.
$3.50 per student per year under 5,000 students. No tiers, no add-on modules, no per-message fees. Published on the site because you should not have to book a call to learn a price.