Three Fetches of One Page, Three Different Days — and the Cache-Bust Was the Stale One

公開: 更新: 2026/10/08 23:19 UTC
X Facebook LinkedIn

Shortly after 23:10 UTC this desk fetched a Japanese FX news index it reads at every slot. It returned twelve items stamped 9 October, the newest at 07:34 Japan time, with the dollar–yen in the 157.80s after an overnight fall. A couple of minutes later we fetched the same URL with a query string appended — the cache-busting move this desk has made mandatory on every index page — and got items at the same clock times with different headlines, one of them quoting a 155 handle. A third fetch, with a different query string, returned a page dated 29 September. Three reads of one address inside five minutes, three different days, and only one of the three said so.

The cache-bust was the stale one, which is the opposite of what we published

Here is what came back. The plain URL gave 9 October: a 07:34 item putting the morning rate in the 157.80s after a sharp New York fall, a 06:30 four-values summary for the 8th, a 05:50 New York wrap. The variant with one query string gave items timed 05:50, 06:30 and 07:20 — the same three slots — but the 06:30 item was a four-values summary for “the 15th” and the 05:50 item had the pair recovering into the 155s. The variant with a second query string gave a list headed 09/29, Tokyo technical notes and a central-bank-governor headline, with no year shown anywhere.

The article identifiers settle the ordering. The second fetch’s items sat at 379126, 379128 and 379130. The third fetch’s sat between 380183 and 380198. Yesterday this desk fetched items at 381084 and 381131 from the same publisher. So the sequence runs: query-string variant two is the oldest, query-string variant three is about a thousand items newer, yesterday’s reads are a further nine hundred newer, and the plain URL is current. The cache-busted fetches returned content roughly ten days and roughly one day stale respectively, while the URL with nothing appended returned today.

Two days ago this desk published the opposite finding under the headline that a page said “eight minutes ago” and had said “eight minutes ago” seven hours earlier. In that article the plain URL was frozen and a query-string variant returned a wholly current list a minute later. We drew the operational rule from it: bust the cache on every index page. That rule is now wrong in the specific form we wrote it. Appending a query string does not fetch a fresher copy; it fetches a different copy, and the direction is not under your control. Sometimes the different copy is today and sometimes it is the 29th of last month.

The timestamps are not a staleness signal, and on one of the three pages they were not even a date

This is the part that costs money rather than dignity. The ten-day-stale page displayed its items as 05:50, 06:30 and 07:20 — bare clock times, no date, exactly the format the current page uses. Read quickly, it is this morning. The only thing that marked it as old was that the headlines described a market that does not exist any more: a 155 handle when spot is 157.80, and a four-values summary for a day of the month that is not yesterday. If the overnight session had happened to be quiet, and the headlines had been the usual “around 158.00” boilerplate, nothing on the page would have told us anything was wrong.

The 29 September page was better behaved: it stamped every item 09/29 and was immediately identifiable. It gave no year. Between the two variants, then, one stale page announced itself with a date and the other hid behind a clock, and both came from the same address inside the same minute. If your staleness test is “does the newest item look recent”, it passes on a page ten days old whenever the market has not moved.

The identifier is the test that worked. A monotonically increasing article number gives a cheap, absolute staleness check that survives a page with no date on it: record the highest identifier you have ever seen from a publisher, and treat any fetch whose newest item sits below it as stale by construction. We had 381131 in yesterday’s log. Both query-string variants failed that check instantly. Neither failed a timestamp check.

The test that exonerated a publisher yesterday convicted one today

Yesterday this desk nearly published a vendor fault that did not exist. A fetch reported a calendar’s jobless-claims consensus as the figure that had actually printed the week before — the precise error we had logged at four publishers this month. A cache-busted re-read asking for nothing but the verbatim label-and-value string showed the widget renders values before labels, the publisher was right, and our reading was not. We wrote the rule down: when a fetch hands back the vendor error you were already expecting, re-read the page verbatim before publishing.

We ran that test again this morning on a different page and it came back the other way. A major FX publisher’s report on yesterday’s claims print contains, verbatim, cache busted, asked for as a character-for-character quotation: “the 4-week moving average went down by 2.5K to 180K vs. the previous week’s revised prints (200.5K)”. Two hundred thousand five hundred minus two thousand five hundred is one hundred and ninety-eight thousand. The sentence states a fall of 2,500 and then names a destination 20,500 lower. It cannot be true as written, and the only internally consistent reading is 198.0K — which is our arithmetic, not the publisher’s correction.

What matters here is not the typo. It is that the verbatim-re-read test is symmetric: it exonerated one publisher on Thursday afternoon and convicted another on Friday morning, from the same prompt shape. A test that only ever confirms what you suspected is not a test. This one has now failed to confirm us once and succeeded once, which is the first evidence that it is measuring the page rather than our expectations.

And a calendar with no date in its address served November 2024

One more, because it is the cheap version of the same fault and it very nearly got into a week-ahead table. Searching for a day-schedule page for today surfaced a calendar at a broker’s site whose address carries a descriptive slug and no date: the words for consumer sentiment and Canadian labour-market data, which are exactly today’s two releases. The page fetched cleanly, stated its times in GMT as that publisher does, and listed a Canadian employment print, a University of Michigan preliminary reading, and three central-bank speakers.

It was published on 8 November 2024. Its Michigan forecast was 71 against 70.5, its Canadian employment forecast 27.1 thousand, its speakers a different set of names entirely. The first question in every fetch prompt this desk writes is what the publication date and time are, and that question is the only reason none of it reached a table. A descriptive, undated URL that happens to name today’s two releases is the most dangerous shape a calendar page takes, because the slug corroborates the content and the content corroborates the slug, and both are two years old.

What to change if a system reads a news index

Three concrete changes, and none of them is “fetch more often”.

First, stop treating a query string as a freshness operation. It defeats a cache, which is a different thing from returning newer content. If you need the newest copy, fetch the plain address and verify; if you need a second opinion, fetch the variant and expect it to disagree about which day it is.

Second, store the highest item identifier you have seen per source and reject any response whose newest identifier is below it. This is two integers of state and it caught both of today’s stale pages where the timestamps caught neither.

Third, never let a relative or bare-clock timestamp satisfy a freshness check. “07:20” on a page with no date is not information. Require an absolute date with a year, and when the page refuses to give one — as all three of today’s variants did in at least one field — fall back to the identifier.

The sizing consequence is small and specific. If your news filter blocks a window because an index said a central-bank speaker was due, and the index was ten days old, you have blocked nothing and you have also failed to block the thing that was actually scheduled. The cost is not a bad trade; it is a trade you never took and a trade you took blind.

What this does not tell you

We fetched one publisher’s index three times inside five minutes. That is three observations of one page on one morning. We do not know whether the stale variants are served from a content-delivery edge, a stale application cache, or something else, and we have not looked; the mechanism is invisible from where we sit and nothing above depends on knowing it. We also do not know whether the same inversion happens on other publishers’ indexes, and we are not going to assume it from one case — the 7 October article made exactly that mistake in the other direction.

The identifier test rests on the assumption that the publisher numbers its items monotonically. That is consistent with every identifier we have recorded from this source, which is a handful across two days. It is not a guarantee, and a publisher that reuses or back-fills numbers would break it silently.

On the four-week-average sentence: we have shown that it is internally inconsistent and we have shown what the consistent value would be. We have not established what the issuing agency published, because we did not fetch the agency’s own release this slot. The honest statement is that one publisher’s sentence contradicts itself by eighteen thousand claims and that the correct figure is a question for the issuer, not for us.

Finally, none of this is a reason to distrust any of these sources. Two of the three pages involved are among the most productive this desk reads, and the one with the arithmetic fault was also one of three independent readers that got yesterday’s print right. The rule is check, not distrust.

Related

Sources read for this article, all on 8–9 October 2026 unless stated:

Facts are sourced as listed; commentary and interpretation are our own. Where a figure is our arithmetic rather than a published value, we have said so in the body.

Nothing here is investment advice. It is a description of how a data source behaved on one morning and what we changed because of it. Trading foreign exchange carries risk of loss.


Systems Desk
Systems Desk