The Corpus·The Source Voice →
Tragedy Academy Studios
Script in, script out.
Every voice model is trained on actors reading copy in a booth. Authenticity Amplified is a seven-year archive of unscripted long-form conversation — the only kind of speech that contains a person figuring something out in real time.
Request dataset access & licensing specs
01 — Where it came from
Seven years of telling the truth out loud.
The Tragedy Academy is a longform conversation podcast about trauma, resilience, and recovery. 178 episodes across seven years, ranked in the top 1.5% of podcasts globally.
The corpus is host audio only. Every clip is a single speaker — Jay Hicks — isolated from his own recordings. No guest audio appears anywhere in the archive. What the conversational setting supplies is the condition: speech produced while actually responding to another person, rather than performed alone to a script.
Every hour was recorded in a treated isolation booth on broadcast microphones — clean acoustics around genuinely messy human behavior. That combination is rare. Most naturalistic speech corpora have poor audio. Most clean corpora have performed speech.
The archive was never built to train a model. It exists because the conversations happened. That's the reason the behavior is in there — nobody was performing.
02 — The problem
Technically excellent, behaviorally empty.
The cause is upstream of the model: a performed read contains no absences. Everything the speaker intended to say arrives, in order, at the pace they planned.
Real speech doesn't work that way. It contains the abandoned clause, the breath before the difficult phrase, the correction mid-sentence, the pause where the thought hadn't resolved yet. Those aren't defects in the signal. They're the observable trace of someone navigating toward something they hadn't reached.
A model trained only on arrivals can't generate a departure. That's the uncanny valley, and no amount of additional synthesis quality closes it.
03 — The asset
What's in the archive.
A continuously growing archive
Seven years of unscripted longform conversation recorded in a treated isolation booth — Shure SM7B, Neumann TLM 103, and Sennheiser MKH 416 through Universal Audio Apollo Twin X. Delivery is by curated, annotated asset rather than by the hour: each volume is a fixed set of verified clips drawn from an archive that keeps growing.
Volume 1 is machine-learning ready
173 curated clips with millisecond-aligned transcripts, tone and emotion descriptions, and structured file paths. The remaining archive is unreleased.
Volume 2 adds the behavioral register layer
200 curated assets, each carrying the social register the speaker was operating in — and, where the register changes mid-clip, the switch point and the observable that makes it audible. Detailed in section 04.
Modular ToneBatch taxonomy
Clips are classified by what the vocalization does — its behavioral function — rather than by which emotion it depicts.
TONEBATCH_04_ReactiveMicro
A dedicated sub-library of isolated backchannels, breaths, sighs, throat clears, and sub-second micro-utterances. Built for latency masking in real-time pipelines.
Single documented origin
One speaker, one source, one owner. Host audio only — no guest voices, no scraped material, no licensed-in or third-party content anywhere in the chain of title.
Expandable by design
New ToneBatches can be created around behavioral weaknesses discovered during model evaluation.
04 — The behavioral layer
Emotion labels describe the surface. This describes the state.
Every corpus on the market is labeled for emotion — happy, sad, angry, neutral. Emotion is what a listener infers. It is not what the speaker was doing.
Volume 2 labels the behavioral register: the social role the speaker occupied while producing the speech. Six of them recur across the archive — host, friend, teacher, comic, vulnerable, troubleshooter — and each is defined by measurement rather than by feel.
Filler density varies by more than twelvefold between adjacent registers — same speaker, same session, same signal chain, minutes apart. One passage runs zero fillers across four hundred and forty words. Another runs four in ninety.
That is not style, and it is not mood. It is a measurable behavioral state, and it is what the model is actually learning when it learns to sound like a person.
The transitions are the asset.
A clip that sits inside one register teaches the model a state. A clip that crosses from one register into another teaches it a change — and change is the thing synthetic speech is worst at.
Clips are treated as spans rather than as containers: an interval with a start and an end, free to overlap other intervals. A span that straddles the boundary between two registers belongs to neither side of it, which is precisely why it is worth the most. Each one carries the switch point timestamped and the observable that makes the switch audible — a rate change, a breath, a pitch reset, a collapse in articulation.
Stated so it can fail
If within-register variance meets or exceeds between-register variance, the register is not real and the label is discarded. The test is applied before delivery, not after. A taxonomy that cannot fail is not a finding.
Closed vocabularies, with a holding lane
Every label comes from a fixed list. Material that does not fit is not discarded — it is logged as unclassified and promoted only after it recurs across multiple independent sources. A pass that returns zero unclassified material is treated as a failed pass.
Early cross-speaker results
The schema has been run against two additional speakers from outside the corpus. Both return a shared core — teacher, friend, vulnerable — plus registers specific to the individual, which suggests the framework describes speakers generally rather than one speaker in particular. Two subjects. Stated as directional, not established.
Why this is not available elsewhere. Emotion labeling is a commodity. Behavioral register indexing requires the framework before the labels are visible at all — you cannot annotate for a state you have no vocabulary to see.
05 — An early result
The model learned a rule nobody labeled.
A production voice model trained on this corpus produces conditional phonetic reduction without ever having been labeled for it. Given the word because, it returns “cuz” in casual clause-linking positions and the full form when the word carries weight.
The inconsistency is the finding. A uniform reduction would just be an accent. Conditional reduction means the model learned the rule from context — a distinction that scripted corpora structurally cannot teach, because a voice actor reading copy says what's written.
06 — Blind listening pilot
Listeners named the difference without being told what it was.
Two synthetic voices rendered the same passage from text — the Authenticity Amplified model and a professional voice model on the same platform. Listeners had no knowledge of the project, the models, or what was being tested, and heard a script written by an independent party specifically to disadvantage the AA voice.
Every response chose the AA model. Their reasons, unprompted:
“because i could here heavy breath”
“felt less continuos, using natural paused and rhythem changes”
“the other speaker sounded more artificial”
“more human, natural, with dynamics and more expressive”
The survey never used the words breath, pause, rhythm, timing, dynamics, natural, or artificial. Listeners supplied every one of them.
Pilot study, n=5. Directional rather than conclusive. Single script, single comparator, fixed presentation order with the comparator first. Expanding.
07 — Demo set
Hear what the corpus produces.
Clips generated from the deployed model, mapped to production contexts. Single-generation renders — no stitching, no comping, no line-by-line assembly.
Each script was written to expose a specific failure mode. If the model can't do the thing, the clip shows it.
![]() | The deployed model Jay Hicks — Synthetic Voice Model Text-to-speech and speech-to-speech models built with Respeecher's Emmy-awarded team and trained on this corpus. Every clip below was generated from it. Listen on Respeecher → |
Games & Dynamic NPCs
JayHicks_Games_NPC_Companion.wav
A companion character responding mid-quest. Contains a self-correction the character does not plan, and a breath that precedes the admission rather than decorating it.
Tests mid-utterance error detection — can the model break, detect, and restart inside a single line?
Take A
Take B
Healthcare & Patient-Facing
JayHicks_Healthcare_Clinical_Presence.wav
A patient check-in that holds a silence rather than filling it, and does not accelerate into reassurance. In high-disclosure clinical contexts, timing is a safety property.
Tests whether a hold reads as consideration rather than a dropout — unbroken room tone through the pause, no gating artifacts, no acceleration into comfort.
Take A — Grounded conversational
Take B — High-trust safety pace
Take C — Warm & intimate
Real-Time Voice Agents
JayHicks_Agent_Realtime.wav
An agent working through a problem while speaking. Audible respiratory preparation, one true self-correction, conditional phonetic reduction.
Tests live computation and memory-retrieval stalls landing in the positions where a person would actually stall. This is the clip used in the blind pilot above.
Animation & Character
JayHicks_Animation_Character_Range.wav
A single character moving through three registers in under a minute — menace, dry humor, weight. Range here is continuous rather than switched between presets.
Tests continuous register change — do the transitions glide, or do they read as preset switching?
Take A
Take B
Documentary & Long-Form
JayHicks_Documentary_Longform.wav
Continuous narration with controlled breath and no prosodic loop. Long-form is the hardest test for synthetic speech — the longer a passage runs, the more chances a listener has to detect repeated cadence.
Tests prosodic looping across an unbroken run. Written with no paragraph resets, deliberately.
Advertising & Brand
JayHicks_Advertising.wav
A brand read delivered conversationally rather than performed. Commercial intent without collapsing into announcer cadence.
Tests whether commercial intent can carry without the read defaulting to announcer posture.
48 kHz · 24-bit WAV · −16 LUFS · −1 dBTP · artifact cleanup only — no EQ, no compression, no gating
08 — Where it applies
Contexts where behavior is the product.
Therapeutic and clinical agents
Systems that can hold a silence correctly instead of rushing to fill it. In high-disclosure contexts, timing is a safety property rather than a polish item.
Real-time conversational agents
Sub-100ms behavioral tokens that consume LLM inference latency as perceived consideration instead of dead air.
Narrative gaming and companions
NPCs with genuine comedic timing, hesitation, and thought pivots — and the ability to distinguish a player's backchannel from an actual interruption.
Gerontechnology and palliative interfaces
Contexts where adoption depends on presence rather than accuracy.
09 — Rights posture
Clean chain of title.
Single speaker, single owner
Host audio only. No guest voices, no scraped material, no licensed-in or third-party content. Every hour was recorded, owned, and curated by one person.
Training-only grant
Rights convey for multi-speaker foundation training, acoustic feature extraction, and disfluency learning.
Likeness retained
Standalone single-speaker voice cloning is excluded absent separate, represented consent.
Documented provenance
Structured for data-lineage disclosure and digital-replica compliance review.
In development
Consented two-speaker conversation
As of September 2026 the guest release carries explicit machine-learning consent, an absolute prohibition on synthetic replication of the guest, and retroactive scope covering prior appearances. Consent is captured at intake as a required, hand-typed affirmation stored with the participant record and its date.
The mechanism is live and is being run back through the archive. Volumes 1 and 2 remain host-audio-only and are unaffected. Consented full-episode conversation — two speakers, natural turn-taking, overlap, repair and entrainment — is a separate forthcoming asset, released only for participants who have affirmatively consented.
Technical documentation. Audio specifications, metadata formats, transcript alignment, taxonomy documentation, register-schema documentation, and sample ToneBatch clips are available on request.
License lived data for your model.
Whether that's the curated Volume 1 corpus, the register-annotated Volume 2, a targeted ToneBatch subset for a specific pipeline problem, staged access to the unreleased archive, or exclusive enterprise rights — let's talk about scope.
Authenticity Amplified is designed as an expandable model-development system. New ToneBatches and new register subsets can be created around specific behavioral weaknesses identified during evaluation, allowing the corpus to evolve alongside the model.
Tragedy Academy Studios · Authenticity Amplified
The Corpus·The Source Voice →
