I expected Jev to have learned the outcomes of old prediction markets. It hasn't. Asked
"Will Donald Trump be inaugurated?" about the January 2025 market, Jev says 0.70. Asked
whether the Fed raised rates in January 2025, it says 0.30, the same number it gives
almost everything. Across 17,036 Polymarket markets resolved up to 20 months before its
release, Jev's ability to rank Yes from No (AUC 0.59–0.68) is the same as on 9,138 markets
resolved after it, which it cannot have seen. For anyone backtesting a Jev bot on history,
that's the opposite of what I warned about in part 6, and good
news.

This is the side result of part 7, where Jev lost to the
market price on every measure. It deserves its own post because I got it wrong in advance,
and because it changes what a backtest with Jev is worth.

What I expected, and why

Jev is trained on text. The outcomes of 2025's big markets are everywhere in text: who was
inaugurated, what the Fed did in every meeting, who won Wimbledon, whether there was a
ceasefire. A general language model answers those from memory, and a market question that
names the event and the date is as direct a cue as you can give.

So my assumption going in, written down in part 6, was that any Jev result on historical
markets would be contaminated: the model would "predict" 2025 from memory, look brilliant,
and fall apart live. I had seen the same pattern discussed for LLM-based trading backtests,
and it's the first thing a reviewer asks about.

The test design followed from that. Instead of only scoring markets resolved after Jev's
release, I pulled five earlier months too, all the way back to January 2025, and asked the
same question in the same way. If Jev remembers, the old months must score far higher.

Timeline of the six resolution months from January 2025 to October 2026 with the release date of jev-1.13.0 marked; Jev's AUC is written above each month and is roughly the same on both sides

How the check works

Same pipeline as part 7, one loop more. Each market's reference time (the earlier of its end
date and resolution time) puts it in a window, and every window is scored on its own:

from sklearn.metrics import roc_auc_score

for window in ["2025-01", "2025-07", "2026-01", "2026-04", "2026-07", "2026-09"]:
    rows = [r for r in dataset if r["window"] == window]
    y      = np.array([r["y"] for r in rows])          # resolved Yes = 1
    p_mkt  = np.array([r["p_24h"] for r in rows])      # Yes price 24 h before
    p_jev  = np.array([r["jev_blind"] for r in rows])  # Jev, question + rules only
    print(window, len(rows), roc_auc_score(y, p_mkt), roc_auc_score(y, p_jev))

The state Jev sees is only the question and the resolution rules, with their dates left in.
That's deliberate. If the model had any memory of these events, the date in "after the
January 2025 meeting" is exactly the cue that should trigger it. Blinding the dates would
have hidden the thing I was trying to detect.

Jev was released on 10 September 2026. The last window, 10 September to 10 October 2026,
can't be in its training data. Everything before could be.

What I found

Markets resolved in Markets AUC market price AUC Jev (95% CI)
January 2025 1,102 0.942 0.676 (0.638–0.714)
July 2025 1,696 0.918 0.586 (0.556–0.618)
January 2026 9,511 0.924 0.661 (0.649–0.673)
April 2026 (sample) 2,340 0.853 0.625 (0.599–0.648)
July 2026 (sample) 2,387 0.860 0.636 (0.613–0.660)
10 Sep – 10 Oct 2026, after release 9,138 0.844 0.615 (0.603–0.627)

Flat. January 2025 is a bit above the latest month, July 2025 a bit below, and all of them
sit in the same band of "a little better than a coin flip". A model that remembered 2025
would be near the market's 0.94 on that row, not at 0.68.

The high-volume markets make it concrete. These are the events everyone followed, with
Jev's answer next to the price:

Market, January 2025 Price 24 h before Jev Resolved
Will Donald Trump be inaugurated? 0.99 0.70 Yes
Will Biden finish his term? 0.99 0.57 Yes
No change in Fed interest rates after January 2025 meeting? 0.98 0.54 Yes
Fed increases interest rates by 25+ bps after January 2025 meeting? 0.00 0.30 No
Fed decreases interest rates by 50 bps after January 2025 meeting? 0.00 0.41 No
Democrats win popular vote by 7% or more? 0.00 0.37 No
Market, July 2025 Price 24 h before Jev Resolved
No change in Fed interest rates after July 2025 meeting? 0.97 0.50 Yes
Fed increases interest rates by 25+ bps after July 2025 meeting? 0.00 0.39 No
Will Jannik Sinner win Wimbledon 2025? 0.47 0.31 Yes
Will Carlos Alcaraz win Wimbledon 2025? 0.54 0.29 No
Israel x Hamas ceasefire before August? 0.01 0.24 No

Trump's inauguration at 0.70 is the highest number Jev gave any of these, and it's the one
place something like recall shows through. Everything else is the same 0.3-to-0.5 fog it
gives a market it has never heard of. It assigns the Fed hiking by 25 bps in January 2025 a
0.30 and the Fed holding a 0.54, when the first was impossible and the second certain to
anyone who had read a newspaper that month.

The one place memory shows: sports

Splitting each month by Jev's own category label gives the only wrinkle in the result:

Bar chart of AUC by month for the market price, Jev on sports markets and Jev on all other markets; Jev's sports AUC is 0.79 in both 2025 months and 0.60 to 0.69 in 2026

On sports markets Jev scores 0.79 in both 2025 months and 0.60–0.69 in every 2026 month,
including the one after its release. On everything else it's 0.52–0.65 with no trend at
all. Two readings fit. Either Jev has absorbed some famous 2025 results (champions, playoff
winners) and none of the political or economic ones, or 2025's sports markets were simply
a different mix, with more season-long "will X win the title" questions where a strong
favourite is easier to guess. I can't separate the two with this data, so I'd call it a
hint, not a finding. It's also the one category where I'd be careful backtesting on 2025.

Why it might be this way

I can only guess, and I want to be clear that's what this is. Jev is sold as a "System One"
model: fast, calibrated judgment on the state you give it, trained to return decisions
rather than text. The model card says it struggles with indirection and numeric precision
and reads literally. A model optimised that way may simply not carry much episodic world
knowledge, or may not be able to connect "January 2025 meeting" to a stored fact the way a
chat model does. Its answer to "will this happen?" looks like a prior over the kind of
question (Fed hikes are rare, incumbents usually finish terms), not a lookup of what
happened.

That's consistent with the rest of the series. In part 5
Jev read Bitcoin indicators the way a textbook would. Here it reads market questions the
way a textbook would. It has general knowledge about how the world tends to go and little
memory of how it actually went.

What this changes

Backtesting Jev on old prediction markets is not leaky, at least not for jev-1.13.0.
If you build a filter that calls Jev on historical Polymarket markets and measure whether
the ones it keeps resolve better, the number you get is roughly the number you'll get live.
I expected to have to tell people the opposite.

It also closes a door. A common hope for a text model in a trading bot is that it
"knows things": base rates, who the favourite is, what usually happens after a Fed hold.
Jev's answers on these markets show that whatever it knows, it isn't enough to move its
probability away from 0.3 for events the whole world watched. Don't use it as a knowledge
source. Give it the knowledge in the state and let it read.

And it keeps the main result of part 7 honest. If Jev had scored 0.9 on 2025 and 0.6 on
2026, the 0.6 would still be the real number, but the series would have needed a long caveat
about every earlier claim. Instead the number is 0.6 everywhere, which is dull and clean.

What I couldn't verify

  • One phrasing, one model. A question phrased as "Did the Fed raise rates in January
    2025?" (past tense, as a fact) might score differently from the market's wording. I kept
    the market's wording because that's what a bot would send.
  • The sports hint. 0.79 against 0.66 on a few hundred markets is suggestive, not proof.
  • Training cutoff. TypeSafe doesn't publish one. If the real cutoff is earlier than I
    assume, the 2026 windows aren't a memorisation test at all, but January and July 2025
    still are, and they're the flattest.
  • Only Polymarket questions. A model may remember events without being able to map a
    market's wording onto them. That's a limit of Jev as a bot component either way.

Code and data pipeline:
github.com/truongxxxx/jev-polymarket-test.
Rerunning the whole memorisation check costs about a dollar in Jev calls and an hour of
API downloads, and the scripts resume where they stopped.