Audit methodology

Translation quality audit: what machine translation does to school messages

Districts translate legally consequential messages with machine translation and nobody has published whether the output is fit for that purpose. This audit runs a fixed corpus of twenty real district message types through five engines, has credentialed bilingual reviewers score them blind against a published rubric, and releases all one hundred outputs in full. The rubric is below; the scores publish in the spring 2027 window.

Last reviewed 2026-08-04 ยท Kastr is pre-launch; we publish dated status rather than logos.

The scoring rubric — five dimensions, scored 0 to 4 blind, with the failure that defines each floor
DimensionWhat is scoredA score of 0 looks likeA score of 4 requires
AccuracyWhether the propositional content survivesA date, a count or a negation changesEvery fact, number and negation intact
School terminologyUS district vocabulary rendered as US districts use it"Principal" rendered as a word meaning "main", "counselor" as a legal adviserTerminology a US district parent would recognise without effort
Register and formalityConsistent usted, appropriate institutional distanceMid-message switch between usted and Consistent throughout, neither cold nor familiar
FluencyReads as written, not as translatedWord order that requires re-reading to parseIndistinguishable from a bilingual staff member's draft
Call to actionWhether the reader still knows what to do and by whenThe deadline or the required action is lost or ambiguousAction, actor and deadline all unambiguous

Reviewers see source and output only. No engine attribution, randomised order, and the same reviewer never scores the same message twice. Inter-rater agreement is reported as Krippendorff's alpha alongside the scores, not omitted when it is inconvenient.

The corpus, and why it is fixed

Twenty message types, chosen because they are what districts actually send and because they span the range from trivial to legally consequential: full closure, two-hour delay, early dismissal, same-day absence, attendance step letter, chronic absenteeism notice, negative lunch balance, immunisation exclusion, IEP meeting invitation, lockdown all-clear, field trip permission, testing window, enrolment verification, head lice notice, bus delay, conference sign-up, medication authorisation, records request response, summer programme enrolment and a general newsletter item.

Each is sourced from a real, publicly posted district communication, lightly de-identified so no district or family is named. Nothing is written by us for the test, because a message written by the people running the test is a message written to be translatable.

The corpus is fixed at edition one and does not change. That is the entire point of a benchmark: re-running the same twenty messages through the same five engines a year later measures engine drift, and a corpus that changes between editions measures nothing at all. Edition two extends the same instrument to Vietnamese, Arabic and Haitian Creole rather than replacing the Spanish set.

Engines, reviewers, and the blinding

Five engines: DeepL, Google Cloud Translation, Amazon Translate, Microsoft Translator and a general-purpose large language model as a baseline, each called through its documented API at its default settings with no per-message tuning. Default settings, because that is how a district's platform actually calls them.

Three credentialed bilingual reviewers — a mix of US district staff who send this correspondence for a living and certified translators — score every output. They see the English source and the Spanish output. They do not see which engine produced it, the order is randomised, and no reviewer knows how many engines are in the test.

Agreement is reported. If three reviewers cannot agree on whether an output is acceptable, that disagreement is more informative than a mean score, and hiding it would be the easiest way to make this audit look more authoritative than it is.

The harmful-error catalogue

A mean score across twenty messages is not what a district needs. What it needs is the list of errors that change what a parent does. Those are catalogued separately, with the source, the output and the consequence stated in plain terms. The classes we expect to find, from the failure modes that recur in school correspondence:

  • Dropped negation. "Do not send your child to school" becoming an instruction to send them. The single most dangerous error class and it is not rare.
  • Shifted deadlines and dates. Date formats reordered, or a relative date like "by Friday" rendered against the wrong week.
  • Institutional terms rendered literally. "Suspension" landing on a word that reads as a holiday. "Free and reduced lunch" as a phrase about discounted food rather than a programme. "Principal" as an adjective.
  • Role confusion. "Counselor" rendered as a legal or financial adviser rather than a school counsellor, which changes who the parent thinks they are being asked to meet.
  • Register collapse. Sliding into in a disciplinary notice, which reads as either condescending or oddly intimate depending on the family.
  • Castilian defaults in a US context. Vocabulary and constructions that are correct Spanish and wrong for a US district audience, which a monolingual administrator has no way to detect.

US-district Spanish, observed rather than prescribed

One output of the audit is a reference list of the terms where US district usage diverges from what a general-purpose engine produces: distrito escolar, consejero escolar, conferencia de padres y maestros, aviso de ausencias, subdirector, ausencia injustificada. It is published as descriptive reference data, drawn from what the corpus and the reviewers show.

It is not a product feature and we want to be exact about that. Kastr has no translation glossary, no terminology-override system and no district term list. If you need enforced terminology today, this audit will tell you what your engine gets wrong and you will have to fix it in the source text before you send.

What Kastr actually does here. We translate with DeepL and we preview. Before a message sends you can render the draft in the languages you are sending it in and read them, which means a bilingual staff member can catch the dropped negation before a family does. The translation cache is keyed on a hash of the source text and holds no index of who a message concerned. We do not publish a supported-language count as a marketing number; DeepL's target set is around thirty languages and the interface exposes fewer than that. Anyone claiming a precise large number is quoting a marketing figure, not a capability.

Questions people actually ask

Is machine translation accurate enough for legally required school notices?

It depends on the notice, and that is the question this audit is designed to answer with evidence rather than assertion. The federal obligation is comprehension: a limited-English-proficient parent must get the same information other parents get, in a form they can understand. For a lunch-menu reminder machine translation is plainly adequate. For a notice that starts a legal clock, a district should be able to explain how it assures the output before it goes out, and human review of the translated draft is currently the only honest answer.

Which translation engine handles US school vocabulary best?

We will publish scores, not an opinion, and the scores are not yet gathered. What we can say now is that engines differ most on exactly the terms that matter in school correspondence — role names, programme names and disciplinary vocabulary — and least on ordinary prose, so a general-purpose quality comparison is close to useless for this use case.

What is the difference between Castilian and US-district Spanish in school letters?

Vocabulary and register, mostly. A general-purpose engine may produce Spanish that is entirely correct and reads as foreign to a US parent, using terms for school roles, programmes and processes that are not the ones their district uses. The failure is invisible to a monolingual administrator, because the output looks like fluent Spanish. That is precisely why the audit publishes every raw output rather than a score.

Which translation errors actually change what a parent does?

Dropped negations, shifted dates and lost calls to action, in that order. An error that makes a sentence read awkwardly costs you credibility. An error that turns 'do not send your child to school' into its opposite is an operational incident. The harmful-error catalogue separates the two rather than averaging them into a single quality figure.

How do I find out which translation engine my current platform uses?

Check the vendor's published subprocessor list, then the data processing addendum attached to your contract; translation providers are almost always disclosed in one or the other because they are processors. If neither names an engine, ask in writing and keep the answer. The audit publishes a vendor-to-engine map from those same public documents, and marks a vendor as unknown where it discloses nothing rather than guessing.

Why do you not claim a specific number of supported languages?

Because a language count is a marketing number and ours would be a small one stated honestly. We translate with DeepL, whose target set is around thirty languages, and our interface exposes fewer. A large round number on a competitor's homepage is worth checking against the engine they actually run, which their own subprocessor list will tell you.

One price. Every feature. Locked for three years.

$3.50 per student per year under 5,000 students. No tiers, no add-on modules, no per-message fees. Published on the site because you should not have to book a call to learn a price.