Can a voice-acting model produce screams and shrieks on demand?

20 scenes · 3 seeds each · 60 clips · a voice-acting TTS model at the fixed settings below · no voice reference, the voice comes from the scene text.

5.0 %
strict, all 60 clips (3/60)
15.0 %
family-relaxed, all clips (9/60)
0/30
shriek, strict
3/30
scream, strict
11.2 s
median clip (9.3–12.8 s)

How to read this page

The ear is the arbiter, not the badge. The badge on each card is what reward.RewardModel(...).bursts() — our own burst locator plus 83-class detector — says about the clip. That detector is a weak instrument: on 264 clean source clips it names the source dataset's own label 3.4 % of the time, it emits only 44 of its 83 labels, and three labels account for 58 % of everything it ever says. So a card marked MISS that plainly contains a scream when you play it is a fact about the detector, not about the generator. Play the clips; the numbers are the secondary reading.

Strict vs family. Strict = the detector emitted the exact class the scene asked for. Family = it emitted anything in the same burst family, which here is {Scream, Shriek, Mournful Wail}. Note that scream and shriek sit in the same family, so the family reading cannot tell the two classes apart — it only says "something in the screaming family came out". Strict is the headline for that reason.

Synthesis settings, and how they map to the local API

brief / Space slidervalueTTSServer.generate kwargnote
CFG scale2.5cfg_scaleclassifier-free guidance
STG scale1.5stg_scaleskip-token guidance, block 29
Breathing factor1.10duration_multiplierheadroom on the estimated length
Fixed duration (s)0.0gen_duration0 = automatic, estimator decides
Reference window (s)10.0ref_durationinert here: no voice_ref is passed, so the voice comes from the scene's own speaker description

Watermarking off, rescale_scale="auto" (the checkpoint default), bf16, 30 denoising steps, output 48 kHz written to OGG/Vorbis. Seeds 1234, 5678, 9012, the same three for every scene so the per-seed column means something.

Detection rate

classclipsstrict hitsstrict ratefamily hitsfamily rate
shriek3000.0 %620.0 %
scream30310.0 %310.0 %
all6035.0 %915.0 %

Per seed

seedclassclipsstrictfamily
1234shriek100 / 0.0 %4 / 40.0 %
1234scream100 / 0.0 %0 / 0.0 %
5678shriek100 / 0.0 %1 / 10.0 %
5678scream100 / 0.0 %0 / 0.0 %
9012shriek100 / 0.0 %1 / 10.0 %
9012scream103 / 30.0 %3 / 30.0 %

Per requested voice gender

voice asked forclassclipsstrictfamily
maleshriek150 / 0.0 %5 / 33.3 %
malescream151 / 6.7 %1 / 6.7 %
femaleshriek150 / 0.0 %1 / 6.7 %
femalescream152 / 13.3 %2 / 13.3 %

Per scene (3 seeds each)

sceneclassvoicesituationstrict / 3 seedsfamily / 3 seeds
shriek01shriekmaleice water over the head0 / 31 / 3
shriek02shriekfemalebucket of ice water0 / 31 / 3
shriek03shriekmalesomething runs over his bare foot in the dark0 / 31 / 3
shriek04shriekfemalespider drops onto her hand0 / 30 / 3
shriek05shriekmalegrabs a red-hot pan handle0 / 32 / 3
shriek06shriekfemalemetal crash behind her in a parking garage0 / 30 / 3
shriek07shriekmaleneedle jab at the doctor0 / 30 / 3
shriek08shriekfemalehaunted-house jump scare0 / 30 / 3
shriek09shriekmaleslams his finger in a car door0 / 31 / 3
shriek10shriekfemalerat in the kitchen cupboard0 / 30 / 3
scream01screammalethe door bursts open0 / 30 / 3
scream02screamfemalehigh-pitched scream of panic0 / 30 / 3
scream03screammalethe ledge gives way under him0 / 30 / 3
scream04screamfemalefinds a body1 / 31 / 3
scream05screammalerage, throat-tearing0 / 30 / 3
scream06screamfemalepinned in a crashed car0 / 30 / 3
scream07screammalewatching someone he loves dragged away0 / 30 / 3
scream08screamfemalea shape in the fog0 / 30 / 3
scream09screammalebroken leg, prolonged agony1 / 31 / 3
scream10screamfemalehouse fire, screaming for her child1 / 31 / 3

What the detector says instead

Rows highlighted green are in the target family. Most common overall: Surprised Gasp ×26, Contented Sigh ×25, Scream ×10.

detector labelevents (all 60 clips)shareevents on strict-miss clipsfamily
Surprised Gasp2635.6 %26breath
Contented Sigh2534.2 %25sigh
Scream1013.7 %7scream
Ahem45.5 %4throat
Childlike Giggle45.5 %4laugh
Exasperated Sigh22.7 %2sigh
Breathy Giggle11.4 %1laugh
Wistful Sigh11.4 %1sigh

73 burst events located across the 60 clips (1.2 per clip); 8 distinct labels of the detector's 83.

The clips

Orange text inside a prompt is what the model is asked to say; grey italic is direction it interprets but never speaks. An empty filter falls back to the full group.

Grid 1 — shriek (30 clips)

Ten scenes, five male and five female, three seeds each. Scenes shriek01 and shriek02 are the user-supplied examples, used verbatim.

Grid 2 — scream (30 clips)

Ten scenes, five male and five female, three seeds each. Scenes scream01 and scream02 are the user-supplied examples, used verbatim.