Script in, script out.

Every voice model is trained on actors reading copy in a booth. Authenticity Amplified is 200+ hours of unscripted long-form conversation — the only kind of speech that contains a person figuring something out in real time.

 

Request dataset access & licensing specs

 


The problem

Most voice AI is technically excellent and behaviorally empty. The cause is upstream of the model: a performed read contains no absences. Everything the speaker intended to say arrives, in order, at the pace they planned.

Real speech doesn't work that way. It contains the abandoned clause, the breath before the difficult phrase, the correction mid-sentence, the pause where the thought hadn't resolved yet. Those aren't defects in the signal. They're the observable trace of someone navigating toward something they hadn't reached.

A model trained only on arrivals can't generate a departure. That's the uncanny valley, and no amount of additional synthesis quality closes it.

 

The asset

  • 200+ source hours. Unscripted longform conversation recorded across years in a treated isolation booth — Shure SM7B, Neumann TLM 103, and Sennheiser MKH 416 through Universal Audio Apollo Twin X.

  • Volume 1 is machine-learning ready. 173 curated clips, roughly 8 hours, with millisecond-aligned transcripts, tone and emotion descriptions, and structured file paths. The remaining archive is unreleased.

  • Modular ToneBatch taxonomy. Clips are classified by what the vocalization does — its behavioral function — rather than by which emotion it depicts.

  • TONEBATCH_04_ReactiveMicro. A dedicated sub-library of isolated backchannels, breaths, sighs, throat clears, and sub-second micro-utterances. Built for latency masking in real-time pipelines.

  • Single documented origin. No scraped, licensed-in, or third-party material anywhere in the chain of title.

 

An early result

A production voice model trained on this corpus produces conditional phonetic reduction without ever having been labeled for it. Given the word because, it returns “cuz” in casual clause-linking positions and the full form when the word carries weight.

The inconsistency is the finding. A uniform reduction would just be an accent. Conditional reduction means the model learned the rule from context — a distinction that scripted corpora structurally cannot teach, because a voice actor reading copy says what's written.

 

Where it applies

  • Therapeutic and clinical agents. Systems that can hold a silence correctly instead of rushing to fill it. In high-disclosure contexts, timing is a safety property rather than a polish item.

  • Real-time conversational agents. Sub-100ms behavioral tokens that consume LLM inference latency as perceived consideration instead of dead air.

  • Narrative gaming and companions. NPCs with genuine comedic timing, hesitation, and thought pivots — and the ability to distinguish a player's backchannel from an actual interruption.

  • Gerontechnology and palliative interfaces. Contexts where adoption depends on presence rather than accuracy.

 

Rights posture

  • Training-only grant. Rights convey for multi-speaker foundation training, acoustic feature extraction, and disfluency learning.

  • Likeness retained. Standalone single-speaker voice cloning is excluded absent separate, represented consent.

  • Documented provenance. Structured for data-lineage disclosure and digital-replica compliance review.

 

Technical documentation

White paper and dataset manifest. For speech researchers, ML engineers, and product leads: audio specifications, metadata formats, transcript alignment, taxonomy documentation, and sample ToneBatch clips are available on request.

 


License lived data for your model.

Whether that's the curated Volume 1 corpus, a targeted ToneBatch subset for a specific pipeline problem, staged access to the unreleased archive, or exclusive enterprise rights — let's talk about scope.

Contact Jay Hicks directly

Tragedy Academy Studios · Authenticity Amplified