All eight demonstrations, the archive test, a typed graph bundle and a rendered map. Both required baselines reproduced exactly. Then two of our conclusions overturned and a hole reported in our own setup.
Both required baselines reproduced exactly: 42 connections in 8 groups at radius 0.30 with the four platform giants alone together, and 94 connections passing the threshold-graph check. Reproducing those is the first scoring criterion, because a run that misses them has a data problem and everything downstream of it is unreliable.
The thresholded geometric construction, at radius 0.35 and a threshold of 130, carrying the confidence figures from the soft-graph run as an annotation on each row. Ship it gated the way the existing stability sections are gated.
The reasoning is that twelve pairings with a stated reason for every company left off is a report section, and forty-two is not. The soft graph does not earn a section of its own, and the run showed why on the data rather than asserting a preference. Because its connection probability falls steadily as distance grows, its confidence ordering is the distance ordering, so on its own it restates a result we already have with extra arithmetic. As an annotation it does real work, because it stops a printed pairing reading as certainty.
The run is candid that the threshold of 130 is not a methodology value. At 140 only two pairings survive, which is too few to print; 130 is the value inside the swept range that reaches a printable length. It says so rather than dressing the number up as a category boundary.
Yes, and the evidence is direct rather than rhetorical.
The rank-only control and the geometric result share 4 connections out of 42. Connection count tracks the composite score at +0.63 in the control and at −0.27 in the geometry, so the two are measuring genuinely different things. Our independent recomputation gives +0.58 and −0.25.
Two specifics carry it better than the correlation does. iQiYi and Netflix sit two points and two places apart in the ranking and are the most differently built of any pair adjacent in it. Amazon and Warner Bros. Discovery are 13 points and 8 places apart, and Amazon is the company Warner Bros. Discovery is built most like.
That is the sentence the weekly cannot currently write. Two companies with the same score are not the same company, and two companies with different scores can be the same kind of company.
The observed number of connections sits 0.45 standard deviations below what distance alone predicts, so the density of the board is entirely ordinary. Across 200 seeds at three settings the null model recovered the platform-giant group zero times. The density is unremarkable and the arrangement is not, and only running the baseline separates those two statements.
All 21 archived weekly states rebuilt cleanly, and the current issue makes 22. The radius was held fixed at 0.30 throughout, so a change across weeks is a change in the board rather than a change in the setting.
ReelShort and DramaBox sit alone together in all 22 weeks without exception. Amazon and Netflix formed a group of their own in all 20 weeks before ByteDance and Warner Bros. Discovery joined the board, and when those two arrived they joined that existing group rather than forming a new one. That is a stronger result than the four appearing together at once would have been.
Week to week, the average company keeps 97.8% of its immediate look-alikes across the 18 transitions in which at least one score actually moved. The worst single week held 87.4%.
What weakens it. Three of the 21 transitions had no score movement and are not evidence of anything. The weeks are not independent, because a score carries forward until an event moves it. The board also grew from 16 companies to 26 across the archive, so counts are not comparable across weeks and only membership is.
We asked for a note on what the run could not do. It returned the most useful document of the set.
It leads with its own failures and lists five judgement calls a reviewer might want to overturn, each with the reasoning attached so the reviewer can disagree on the data rather than on authority. Four things in it we would not have found by reading the results:
The hole was real and the ambiguity was ours. We resolved it in the opposite direction from the obvious one. The rubric is now explicitly open and agents are told to read it, and only the sealed answer stays sealed. An agent that has to guess the criteria is being tested on mind-reading.
Three things the run names and does not close. Whether the peer groups change when distance reflects the index's own weights rather than treating the five dimensions equally, which the run calls its most obvious next test. Whether a geometric section repeats what the movers section already says, which is the rule that decides whether it can ship. And whether the second unoccupied position it reported survives the sensitivity check it ran on the first.
| Construction | Verdict | Reason |
|---|---|---|
| Thresholded geometric | Ship | Twelve pairings is a report section. Forty-two is not. |
| Soft geometric | Annotate | Its confidence figures belong on the shipped table, not in a section of their own. |
| Waxman null model | Run, never print | The check behind the printed claim. |
| Geographical threshold | Drop | On this board a high score and a distinctive position arrive together, so strength never overcomes difference. |
None of it ships yet. One system has run the brief, no reader has been shown the output, and the recommendation rests on the properties of the results rather than on anyone having read them. DeepSeek and Kimi run the same file next, unedited.