Yesterday this desk published two live disagreements about what the US flash surveys were forecast to do: services at 55.8 or 56.0, manufacturing at 53.6 or 53.5. Both printed at 13:45 UTC. Services came in at 58.7 and manufacturing at 57.0, which means the two-tenths gap was swamped by a 2.7-point beat and the one-tenth gap was swamped by a 3.4-point beat. Both disputes were irrelevant. That was the prediction, and it held. The problem is what we found in the next column over.
Mark the splits: both dead, and by a factor of thirteen and thirty-four
Services printed 58.7 against a consensus of 56.0 on one reader and 55.8 on another. The beat is 2.7 or 2.9 points — thirteen and a half times the width of the disagreement. Manufacturing printed 57.0 against 53.6 and 53.5. The beat is 3.4 or 3.5 points, thirty-four times the gap. The composite came in at 58.4, the strongest in about five years. S&P Global’s own chief business economist reads the set as consistent with annualised growth around 5 percent.
Yesterday this desk published a rule off the euro area board: a consensus gap matters when the print lands inside it, and not when the gap is wide. That was built on one morning’s data, where a one-tenth gap changed a sizing decision and a seven-tenths gap did not. Today is the first out-of-sample test and the rule survives it cleanly — two gaps, both narrow, both irrelevant, because both prints landed miles outside them.
We are not upgrading it. Two mornings is two mornings. What we will say is that it has now been right about four disagreements in a row, and that it is cheap to apply: before you widen a stop because your feeds disagree, check whether the plausible print range actually intersects the gap. Most of the time it does not.
The column that was actually wrong was the previous one
Here is what we did not check yesterday. Every calendar we read carried the August values as 56.5 for services and 53.9 for manufacturing. The publisher reporting the release carried them as 56.8 and 53.2. Those are not typographical errors and neither party is wrong. The calendars are quoting the August final. The release is quoting the August flash.
Trading Economics settles it explicitly on both series: the August services figure was revised lower from a flash estimate of 56.8 to a final of 56.5, and the August manufacturing figure was revised higher from a flash of 53.2 to a final of 53.9. S&P Global’s own August release, read directly, carries 56.8 and 53.2 in its headline table. So both numbers in circulation are the issuer’s, from two different documents a month apart.
This is the same class of defect this desk has been publishing all week — a field whose identity is not what its label says — except that it sits in the one column nobody argues about. Traders spend the morning disputing the forecast. The prior is treated as settled fact. It is not.
The two revisions went in opposite directions, which is the part that bites
If every flash were revised the same way, the ambiguity would be a constant and it would wash out of any month-on-month calculation. It does not. Services was marked down three tenths; manufacturing was marked up seven tenths. So the change you report for September depends on which prior you keyed to, and it depends differently for each series.
- Services: 58.7 against the final 56.5 is a gain of 2.2 points. Against the flash 56.8 it is 1.9.
- Manufacturing: 57.0 against the final 53.9 is a gain of 3.1 points. Against the flash 53.2 it is 3.8.
Seven tenths of a point on the manufacturing month-change, from a choice of source document that no vendor labels. If your system scales position size by the magnitude of the month-on-month move rather than by the surprise against consensus — and plenty do, because the surprise is the thing everyone knows is disputed — then that parameter just moved by nearly a quarter of its own value without anybody publishing a correction.
The methodologically correct comparison for a flash release is flash to flash: same sample coverage, same point in the collection cycle. That is the comparison the issuer makes in its own commentary, and it is the one your calendar is not showing you.
The issuer contradicts itself too, and that is the confirmation
We wanted a check on the mechanism rather than on the numbers, and the issuer supplies one. S&P Global’s July flash release gives July manufacturing at 53.8. Its August flash release, in the comparison column for the same month, gives July at 53.9. One tenth, same publisher, same series, two documents.
That is not an error. It is the flash being superseded by the final between the two publications, exactly as described above, and it means the pattern is not an artefact of how any particular vendor copies numbers. It is structural to the release. Any series you build by pasting together the “previous” fields from a flash calendar is a mixture of two different vintages, and the mixture is not documented anywhere in the feed.
What moved, and what it cost
The surveys were not a quiet beat. Employment growth in the survey was the strongest since June 2022, and input cost inflation was the highest in about four years, which the compiler attributes to fuel and transport. The reaction was a straight rate-repricing: yields to new fifty-two-week highs across the curve, with the five-year briefly through 5 percent for the first time since 2007 and the ten-year closing near 5.11. October Fed hike odds are reported at 64 percent on one feed and around 70 on another.
We are publishing both probability figures rather than picking, on the same principle as everything above. A six-point spread on a binary event a month out is not noise if your sizing is conditional on it.
What this does not tell you
It does not tell you the calendars are wrong to print the final. The final is the better estimate of what August actually was. It is simply not the number the September flash is being compared against in the issuer’s own commentary, and no feed we read says which one it is showing.
It does not give you a revision rule. Two series, one month, one revised down and one revised up: that is not a sample. We have not gone back through the history to measure whether flash-to-final revisions are biased in either direction on either series, and until someone does, do not build a correction factor from this article.
It does not settle the September figures either. Today’s 58.7 and 57.0 are themselves flash estimates and will be revised when the final lands. Whatever you conclude from them, conclude it knowing that the numbers are provisional by construction.
And it does not tell you the beat was real in the economy rather than in the survey. A diffusion index measures the share of firms reporting improvement, not the size of it. Five percent annualised growth is the compiler’s reading of its own index, and it is a reading, not a measurement.