# CHRIS'S PROMPT BIBLE — living document
*Dictated by Chris, encoded by Claude. Every shots.json generator must obey this file.*

## LAW 0 — THE INHERENCE TEST (the master rule)
For every candidate detail, ask: **will the model infer this correctly from
what the prompt already establishes?**
- **Inferred correctly (inherent):** SAY NOTHING. Naming an inherent effect
  amplifies it — the model treats mention as emphasis. (Winter dusk already
  makes breath visible; naming it makes a chimney.)
- **Not inferred (non-inherent):** SPECIFY, or the model's statistical default
  flattens it. (Real traffic's heterogeneity is inherent to reality but NOT to
  the model — unspecified, it flocks.)

The prompt is a customs border: only non-inherent details get through.
Laws 1 and 2 below are the two big applications of this test.

## LAW 1 — THE FLOCK LAW (density homogenizes)
Any element appearing in multiplicity will be rendered as near-clones: same
model, same plate, same spacing, same state, defects at regular intervals.

**Countermeasures — specify DISTRIBUTIONS, never let the model default:**
- **Identity variance:** name the mix explicitly (vans, cabs, a bus, sedans,
  SUVs; many brands, colors, ages; some dirty, one damaged). Applies equally to
  crowds (heights, builds, gaits, clothing eras), windows (lit/dark/curtained),
  foliage, birds, rain streaks, wear and tear.
- **Spacing variance:** state it numerically — "roughly 30% variance in the
  gaps," "unevenly spaced."
- **State variance:** enumerate the states and distribute them — headlights:
  some full beams, some bluish LEDs, several daytime running lights, a few
  parking lights only, one off. License plates: different states, mostly dirty.
- **Random defects:** broken/missing items must be "at random positions, not
  evenly spaced" — never "1 in 10" unqualified, or the model places every 10th.
- **The Outlier Rule:** every dense field gets at least one deliberate
  grid-breaker — a driver half-nosed into the next lane craning to see ahead.
  The outlier is what makes the shot read as witnessed, not generated.

## LAW 2 — THE BAKED-IN LAW (mention = amplification)
Context already generates certain effects. Winter evening → visible breath.
City dusk → some steam, some exhaust. Wet street → reflections. Restating a
baked-in effect DOUBLES it: "smoking manhole" → volcano; "visible breath" →
cigar plume.

**Countermeasures:**
- Inventory what the stated time/season/weather already produces; do NOT
  restate those effects.
- When an effect must be mentioned, dampen it explicitly: "faintly visible
  breath," "a thin wisp," "barely," "occasional."
- Intensity adjectives are multipliers — choose the mild register for anything
  meant to be ambient.

## LAW 3 — DURATION DISCIPLINE (seconds are a budget, not a prize)
Legacy bias from the 4-second era: when 8–10s became available, the instinct
was to milk every second. Wrong. Extra seconds don't automatically add value —
they add drift risk, morph risk, and credit burn, while the value lives in the
CUT.

**Countermeasures:**
- Choose duration by the shot's narrative function, not by the maximum
  available: a glance/insert = 4s; an establish or mood beat = 6–8s; a
  developing action that genuinely evolves = 8–10s.
- If the shot's content doesn't CHANGE meaningfully across the extra seconds,
  cut the duration — the edit will thank you.
- More seconds in one clip = fewer total shots per credit pool. Coverage beats
  length.

## LAW 4 — THE CONSISTENCY STACK (Flow: Characters & Ingredients)
Characters can be objects; objects can be characters. Any entity that must
stay identical across shots — a face, a jar, a notebook — becomes a Character
or Ingredient ONCE, then gets referenced, never re-described. (Re-describing =
re-rolling the dice on identity; the Young MC lowrider lesson.)
- Totem objects (the tip jar) get the same treatment as actors.
- Build stills with the HIGHEST image model (Nano Banana Pro tier) — the extra
  tokens are trivial, and the photorealism difference shows on large screens.
- Never generate 20 of the same thing. One deliberate generation, judged, then
  re-rolled only with intent.

## LAW 5 — THE NO-LUCK DOCTRINE
Do the work: enumerate the specifics yourself, every time. Never hand a
creative choice to the model's dice ("add personality" is a prayer; "masking
tape hand-lettered TIPS with a doodled crown, a peeling dancing-bear sticker,
a bead bracelet around the neck" is a specification). And when a loose prompt
DOES land — that's luck, not method. A stochastic machine doesn't owe you the
same result twice. Past success without specification is not evidence;
it's a warning you got away with one.

## LAW 6b — SITUATIONAL LANGUAGE (Chris: "tunnels don't have mouths")
Write locations the way a location scout speaks, not the way a novelist does.
Anatomy/poetry metaphors ("tunnel mouth", "the city's throat") are AI-speak to
the model and produce invented geography. Use situationally aware, real-world
terms: "the eastbound entrance of the Lincoln Tunnel at [street]/[street]",
"the exit ramp onto Dyer Avenue", "the curb on the approach". Poetry lives in
the VO and the edit — never in the render instruction.

## GROUND TRUTH — FROM TRUDI HERSELF (phone call, 2026-07-30)
**THE NAME PROTOCOL:** She does not like being called Queen. THE TIP JAR QUEEN
is the film's title, not her nickname. In every prompt, dialogue line, and
character card she is **Trudi** (or a variation she uses). The driver's line
becomes "Mornin', Trudi." Flow character to be renamed TheQueen → Trudi.
No character in the film calls her Queen.

**THE LOCATION (canonical, replaces all "tunnel mouth" language):**
Her spot: **428 W. 36th Street — the Dyer Avenue corridor (34th–36th leg) at
the Lincoln Tunnel approach.** What the camera actually sees there (from her
own Street View captures): the rail overbridge crossing the trench; concrete
trench walls with yellow chevron markings; a YIELD sign where the ramp
surfaces; two lanes per portal, never more; white bollards and a corner
planter; the crosswalk; Hudson Yards towers stacked behind. Buses in this
corridor are real (Academy and NJ Transit coaches crawl right past her spot).
Her old camp corner (Google Street View, Oct 2019): blue tarp against the
building wall beneath a billboard, roll-down gate, corner traffic light.

**HER SIGNS ARE ART:** Real signs are hand-drawn cardboard with ILLUSTRATED
bubble lettering, colored in purple/blue/green marker — poster-craft, not
scrawl. Register of her written voice: courteous and warm ("PLEASE, I'M
KINDLY ASKING FOR A LITTLE HELP! THANK YOU! GOD BLESS"). The film's TIPS
sign, any cardboard signage, and her lettering in inserts must echo THIS
style. (Production idea, pending: commission the real Trudi to hand-letter
the film's main title card — paid, credited.)

**AGE REGISTER:** The real Trudi reads younger than the "late 50s" I wrote
into early prompts. Soften age descriptors; identity comes from the character
reference, not from age adjectives (Law 4 anyway).

## LAW 7 — DIALOGUE ATTRIBUTION (lines bind to the strongest presence)
E-report, 2026-07-30: "she says 'Mornin', Queen' to the car." A line given to a
barely-established speaker (an off-frame arm) migrates to the most salient
character in the shot — the anchored @character wins custody of any loose line.
**Countermeasures:**
- ESTABLISH THE SPEAKER ON CAMERA BEFORE THE LINE (synergy with L8): "In the
  nearest car, a driver leans toward the open window, looks at her, and says…"
- Speaker-first sentence order; the line comes AFTER the speaker exists.
- Give the non-speaking featured character a mouth-free occupation ("keeps
  writing, nods") — negations ("she doesn't speak") are Law-2 traps.
- OR pre-assign the line off-screen (Chris): "a voice heard off camera says
  'Mornin', Queen'" — explicit O.S. attribution breaks the salience pull
  without needing the speaker in frame. Either fix is acceptable; choose by
  whether the shot WANTS the speaker's face.

## REVIEW CODES (Chris's shorthand)
- **E:** = error/defect report — respond with diagnosis + fix, never thanks.

## LAW 6 — THE GROUND-TRUTH LAW
When a scene depicts a real place and time, don't invent a plausible
distribution — PULL THE TRUE ONE (toll classifications, registration data,
transit schedules, sales rankings). Plausible-sounding is where hallucination
hides: "yellow cabs at the 6 AM inbound Lincoln Tunnel" sounds like New York
and is mathematically near-zero.

**THE 6 AM EASTBOUND LINCOLN TUNNEL FLEET SPEC (data-grounded, reusable):**
Roughly two-thirds private cars, dominated by compact crossovers — Honda CR-V,
Toyota RAV4, Nissan Rogue, Chevy Equinox, Jeep Grand Cherokee — several Tesla
Model Y, some Honda Civic and Toyota Camry sedans; colors mostly white, black,
gray, silver, an occasional red or blue. About one vehicle in five is a BUS:
white NJ Transit and private commuter coaches, plus a jitney shuttle van.
Work vehicles mixed through: a contractor van with ladder racks, a pickup with
toolboxes, one box delivery truck. A couple of black livery SUVs. NO yellow
cabs. License plates predominantly NJ cream-colored, a few NY, mostly dirty.

## THE CHARACTER-SHEET FORMULA (Chris, verbatim — the canonical creation prompt)
> "Full-body triptych, three distinct views: front facing, 3/4 side view, and
> back view. High resolution, flat studio lighting, consistent anatomical
> proportions across all views, solid white background. Same body, same outfit."

Characters are built from a TURNAROUND, not a portrait: three angles of one
body teach the model the person's geometry instead of one lucky viewpoint.
Append the character's identity description to this template, generate on the
highest image model, and THAT image (or pair) becomes the character's
appearance reference. A cinematic scene still is never a character foundation.

## THE FLOW LANGUAGE (official grammar — learned 2026-07-30, no more winging)
Sources: Google Flow Help "Create videos" + "Manage projects, assets & collections."

1. **`@` is the reference operator.** Type `@` in the prompt box to search and
   select saved assets BY NAME. `@CharacterName` for characters, `@me` for your
   avatar, `@Voice:Name` for voices. The @ mention attaches the saved identity.
2. **Ingredients: max 3 per prompt.** Add via the + button (project assets or
   upload), drag-in, or @ mention. Best source images: subject on a PLAIN or
   segmented background (our cinematic stills work but plain-bg versions are
   the textbook input).
3. **The prompt assigns ROLES, it never re-describes identity.** Official
   example: "The woman, whose torso is the lava lamp, walks down the foggy
   street" — the ingredient carries WHO/WHAT, the prompt carries HOW it
   functions in the clip. Docs verbatim: "Ensure prompts complement visual
   inputs rather than contradict them." Re-describing a saved asset = fighting
   your own reference = random re-roll (the tip jar mistake).
4. **Characters are built in the Characters tab:** detailed description +
   1–2 appearance images (≥1 required) + a NAME (+ optional custom voice).
   Then `@Name` forever after. Voice references only work in ingredient-based
   generations.
5. **Frames mode:** start-frame/end-frame slots animate stills and build
   transitions — the prompt describes the transition, not the image contents.
6. **Name assets cleanly** (TipJar, TheQueen) so @-search finds them; assets
   can be renamed and organized into collections.

## THE AUDIO GRAMMAR (Omni Flash native sound — no lipsync, no post)
Omni Flash generates synchronized dialogue, effects and ambience in-shot.
- **Dialogue: COLON FORMAT, NEVER QUOTATION MARKS** (ratified 2026-07-30 on
  research: quotes trigger burned-in subtitles ~40% of the time). Write:
  `The driver says warmly: Mornin', Trudi.` — speaker, verb, colon, bare line.
- **Lipsync degrades after ~6–7 seconds** — land dialogue lines in the first
  half of a shot, never the back half of anything over 8s.
- **Ambience:** specify a short sound bed ("layered idling engines, one distant
  horn") — same inherence test applies: don't name sounds the scene already
  implies unless dampening or featuring them.
- **Always append a caption guard:** "no subtitles, no captions, no on-screen
  text" — Veo-family models love burning captions into dialogue shots.
- **Put speakers ON CAMERA (Chris, L8):** don't stage dialogue off-screen out
  of lip-sync fear — that's legacy caution from older models. Omni Flash
  renders visible speaking faces that look and sound real; on-camera speech
  adds a whole layer of life. Hide a speaker only as a deliberate craft choice
  (mystery, reveal), never as a hedge.
- **Voice consistency:** Characters can carry a custom voice; reference with
  `@Voice:Name` — voice references only work in ingredient-based generations.
- Character dialogue + native audio means whole dialogue scenes are single
  generations. Score and VO still come from our own pipeline in the edit.

## THE FRAMES-FIRST PIPELINE (ratified 2026-07-30)
Practitioner consensus, adopted as doctrine: **author every shot's start frame
as a Nano Banana Pro still, then generate frames-to-video.** "The model
follows the image much more closely than it follows your adjectives."
- The stills pass IS the storyboard: cheap, reviewable in one scroll, hearted
  or killed by the director before any video credit burns.
- Fixes at the root: static-shot unreliability, arc/orbit failures, character
  drift, geography invention.
- Video prompts on frame-based shots describe MOTION and SOUND only — the
  image already carries identity, composition, and place.
- Never use Flow's Extend (silently downgrades to Veo Lite; characters age
  across extensions). Hold last frames in the edit, or split shots.
- No seeds exist in Flow; nothing may depend on them.

## STANDING ORDER
Apply both laws to ALL dense or context-implied elements — traffic, crowds,
lights, weather, decay, wildlife, texture — not just the category the lesson
came from.

---
*Lesson log: L1 2026-07-30 — the tunnel traffic flock + the smoking manhole.*
