Skip to content
HOLLYWOOD METRICS
HomeAboutCraftNewsAnalyticsScriptsBooksPricingAPIMCP
The BookThe NullInsidePrinting & PostBasket
A single taut horizontal wire over a dark grid, lit from one side
Books/The Null

THE NULL

Twenty measurements of a screenplay, six measures of what the film did, and the whole matrix printed rather than described.

THE RULE BEHIND IT

A correlation runs from minus one to plus one. Minus one means two things move exactly opposite; plus one means they move exactly together; zero means knowing one tells you nothing about the other. Square it and you get the share of the variation in one that the other accounts for.

That squaring is what makes the numbers below smaller than they look. The strongest correlation in the whole table with the composite reception score is r = 0.068, and 0.068 squared is 0.0046. Four tenths of one per cent.

There is no threshold at which a correlation becomes real. What there is, is a comparison: the production budget alone correlates with worldwide gross at about 0.64 on the same catalogue, which is what a relationship of substance looks like next to these.

A machinist's spirit level on a dark slab, its bubble dead centre
A correlation of zero is not a small reading. It is the instrument sitting level.
Twenty screenplay measurements against how the finished film was receivedOn a correlation axis running from minus one to plus one, all twenty measurements of the screenplay text fall between -0.068 and 0.059. The production budget alone correlates at 0.64.THE FULL SCALE A CORRELATION CAN OCCUPY-1-0.50+0.5+1the budget, aloneall twenty togethertwenty measurements of the screenplay, hereTHE SAME TWENTY, MAGNIFIED TEN TIMES-0.10-0.050.00+0.05+0.10strongest of the twenty: vocabulary richness, -0.068
Pearson correlation between each of twenty structural measurements of a screenplay and a composite score for how the finished film was received, over 1,571produced films. The two gold marks are the square roots of the R-squared of two models of the same catalogue: the production budget alone, and all twenty screenplay features together. Sources: the book’s fact base, keys screenplay.correlation_matrix_shipped, models.r2_budget_only and models.boxoffice_model_r2.
r = 0.068
Of
1,499 screenplays re-measured with the corrected parser and joined to their outcomes
How
the largest absolute correlation between any rebuilt feature and the composite quality score
But
The instrument was rebuilt, every structural measure changed, and the answer did not. That is what makes the null result a finding rather than a failure.
41%
Of
11,427 films with a reported budget of at least $10,000 and a reported worldwide gross of at least $1,000
How
R-squared of an ordinary least squares fit predicting log worldwide gross from log production budget alone
But
One number, available before anyone reads the script, and it is the benchmark every cleverer model has to beat.

THE WHOLE MATRIX

Every one of the 20 measurements against every one of the 6 outcomes, over 1,571 produced films. Shading is keyed to a correlation of 0.2, not to 1: keyed to the full range this table would be one flat colour, which is true and tells you nothing.

Pearson correlation between twenty screenplay measurements and six measures of film reception
MeasurementMaster scoreAudience ratingCritic scoreWorldwide grossReturn on budgetAcademy Award wins
Scene countscene_count-0.040-0.001-0.045+0.029-0.021+0.000
Average scene lengthavg_scene_length+0.059+0.054+0.035+0.019+0.002+0.020
Interior to exterior ratioint_ext_ratio-0.022-0.040-0.025-0.068-0.028-0.022
Scene length variancescene_length_variance-0.018-0.023-0.001-0.004-0.017-0.014
Total pagestotal_pages+0.039+0.087+0.022+0.142-0.063+0.108
Transition densitytransition_density-0.033-0.024-0.003-0.055+0.011-0.004
Distinct speaking charactersunique_character_count+0.009+0.048+0.015-0.053+0.021+0.011
Dialogue ratiodialogue_ratio+0.002+0.042+0.028-0.055+0.007+0.005
Top three character dominancetop3_character_dominance-0.019+0.024+0.015-0.045-0.039-0.007
Average speech lengthavg_dialogue_length-0.014+0.020+0.002-0.058-0.013-0.017
Character introduction ratecharacter_intro_rate+0.007+0.029+0.009-0.055+0.068-0.006
Vocabulary richnessvocabulary_richness-0.068-0.087-0.134+0.003-0.032-0.059
Average word lengthavg_word_length-0.049-0.068-0.088+0.044-0.021-0.033
Sentiment, meansentiment_mean+0.057+0.091+0.144-0.007+0.001+0.043
Sentiment, variancesentiment_variance+0.002-0.032-0.024+0.052+0.032-0.010
Sentiment arc slopesentiment_arc_slope+0.022+0.044+0.011-0.024-0.015-0.022
Action ratioaction_ratio+0.006-0.035-0.029+0.061-0.004-0.009
Capitals densitycaps_density-0.013-0.013-0.033+0.160-0.051-0.020
Exclamation densityexclamation_density+0.026+0.014+0.044+0.157+0.019+0.020
Question densityquestion_density+0.036+0.059+0.122-0.096+0.033-0.054

graded screenplays, correlating each structural feature against each outcome measure. Pearson correlation per pair, over rows where both are present. Fact base key screenplay.correlation_matrix_shipped. The whole catalogue is available as JSON.

THE FIVE LARGEST VALUES IN THE TABLE

Named, because they are the strongest thing twenty measurements of a screenplay ever found, and because a page that only says “nothing here” is hiding its own best evidence against itself.

  1. +0.160Capitals density against worldwide grosssmall — and this is the biggest one
  2. +0.157Exclamation density against worldwide grosssmall
  3. +0.144Sentiment, mean against critic scoresmall
  4. +0.142Total pages against worldwide grosssmall
  5. -0.134Vocabulary richness against critic scoresmall

The largest is capitals density against worldwide gross, and it is almost certainly not about storytelling: capitals density counts words of two or more capital letters, which means it counts scene headings and character cues as well as emphasis, so it partly measures how densely formatted a file is. That is the kind of thing a table like this is for.

An empty reading chair and side table under one lamp, in darkness

The strongest thing eighteen months of measurement found explains four tenths of one per cent of what happened next.

Chapter fourteen

THE QUESTIONS THIS TABLE GETS ASKED

— Does this book say screenwriting does not matter?

No. It says that twenty structural measurements of a screenplay text - scene count, dialogue ratio, sentiment arc and seventeen others - carry almost no information about how the finished film was received, across 1,571 produced films. The strongest correlation found was 0.068, which explains under half a per cent of the variance. That is a finding about what those measurements can see, not about whether writing matters. Chapter fourteen states the claim precisely and chapter seventeen states what survives it.

— Is the analysis reproducible?

Every number in the book resolves to an entry in a machine-readable fact base carrying the value, the sample size, what that sample is a sample of, the source file, the method and the caveat. The manuscript is checked against it before every build: a number without a citation stops the build, and a table that disagrees with the fact it cites stops the build. Appendix A sets out the method in full.

— Why are so many films missing from the money figures?

Because they never reported. Only 15.46 per cent of the 74,571 films in the catalogue carry both a production budget and a worldwide gross, and the ones that do are not a random sample of the ones that do not. Every profitability statement in this book, and in every book like it, describes that fraction. Part One computes each of them twice - once over the films that reported, once over a bound that treats every non-reporting film as a failure - and prints both.

— What does the book get wrong?

Three findings this project published before the book existed are withdrawn in chapter sixteen, because the parser that produced them was reading three quarters of the corpus as having no dialogue. Chapter fifteen is the bug, chapter seventeen is the rebuild, and chapter nineteen is what happened between the fix being written and the numbers being corrected. It is the most uncomfortable chapter in the book.

— Can I check a number without buying the book?

Yes. The correlation matrix that carries the central finding is printed in full on this site, with the sample size and the method beside it, and so is every headline figure. The reference tables in Part Four are the part that needs 200 printed pages.

— Do you sell through Amazon as well?

Yes, and the paperback costs less here. A retailer takes a cut of every copy; buying direct removes it, and the ebook comes with the paperback at no extra charge, which no retailer can offer.

— How long does a copy take to arrive?

Each copy is printed when it is ordered, so the window covers manufacturing as well as carriage. Standard post to a United States address is 11 to 15 days; to the United Kingdom, 7 to 9. The delivery page carries every zone, and checkout quotes the real services for the country you name.

— What is in the reference section?

About a thousand rows: every decade and every genre on one template, the hundred highest master scores, the fifty largest returns and the fifty largest grosses, the highest-scoring film of each of 107 years, and the twenty largest gaps between a budget and a gross. Each table prints the filter that produced it, which is the part most published lists leave out.

The argument above is chapters fourteen to eighteen. What the rest of the book adds is how this number was got wrong first, what it cost to find that out, and the thousand rows of reference tables the two hundred printed pages exist for.

The paperback, direct — 19.99 USD
HOLLYWOOD METRICS

The definitive mathematical oracle for cinematic success and creative script intelligence.

74,000+ Films100 Years1.1M+ Reviews
PRODUCT
  • Dashboard
  • Pricing
  • Script Analysis
  • Watch the Demo
  • The Red List
  • News
  • Analytics
  • API
  • MCP Connector
COMPANY
  • About
  • Track Record
  • Trust & Script Privacy
  • Privacy Policy
  • Terms of Service
CONNECT
  • ▶YouTube
  • ◉Instagram
  • ♪TikTok
  • @Threads
  • fFacebook
Get Started Free →
© 2026 Hollywood Metrics. All rights reserved.Powered by bespoke agents and autonomous agentic workflows