Can Jev beat Polymarket? I asked it about 26,174 resolved markets
Part 7 of 8 · Jev in trading, tested
Can Jev beat the Polymarket price? Not from the question text, no. I sent Jev the question
and resolution rules of 26,174 resolved Polymarket markets and asked one thing: will this
resolve Yes? On the 9,138 markets resolved after Jev's release, so nothing it could have
seen in training, it scored an AUC of 0.615. The Yes price traded a day before scored 0.844.
Telling Jev the price made it worse than the price. Blending the two added nothing. And Jev
answered "more likely Yes than No" on 2.4% of the markets, because it mostly says 0.3 to
everything.
This part is the main result, with the code. The side result I didn't expect, that Jev
doesn't remember old market outcomes either, gets its own write-up in
part 8.
Why this test
In part 6 I wrote that Jev had failed every price-only test I'd
run, but that I hadn't tested a task where the information isn't in the price. Prediction
markets are the obvious candidate. Each market is a sentence in plain English plus a page
of resolution rules, which is exactly the kind of input Jev is built for. And "use Jev to
find mispriced Polymarket markets" is a popular pitch right now: scan thousands of markets
in code, let Jev judge which ones are worth a trade, size and execute in code.
Polymarket also publishes everything I need to check that pitch: every resolved market, its
text, and its full price history. Nobody paid me for this and I don't sell anything built
on Jev.
How I tested it

| Setting | Value |
|---|---|
| Markets | Resolved Yes/No markets with at least $10k volume, from Polymarket's public Gamma API |
| Windows | Markets resolved in Jan 2025, Jul 2025, Jan 2026, Apr 2026, Jul 2026 (samples) and 10 Sep – 10 Oct 2026 (all 9,138) |
| Baseline | Last traded Yes price at least 24 hours before the earlier of end date and resolution time |
| Jev input, blind | {market_question, resolution_rules}, rules cut at 1,200 characters |
| Jev input, priced | The same plus the 24-hour-earlier Yes price, labelled as the traders' implied probability |
| Jev questions | One Noul: "this market resolves Yes"; one Choice: category |
| Model | jev-1.13.0, released 2026-09-10 |
| Cost | 52,348 requests, 722 input tokens each, about $1.60 in total; median latency 284 ms |
The 24-hour price is a hard baseline on purpose. A bot that wants to beat the market has to
beat the market's current price, not a coin flip. The question is whether Jev brings anything
the price doesn't already contain.
Getting the markets and the price
The Gamma API returns closed markets with their text and final outcome. I keep the ones
UMA actually resolved, with a clean 0/1 outcome and a Yes/No pair:
url = ("https://gamma-api.polymarket.com/markets?closed=true&limit=100"
f"&offset={off}&end_date_min={day}T00:00:00Z&end_date_max={next_day}T00:00:00Z"
f"&volume_num_min={min_volume}")
for m in get(url):
prices = json.loads(m["outcomePrices"]) # e.g. ["0", "1"]
if m["umaResolutionStatus"] != "resolved": continue
if set(prices) != {"0", "1"}: continue
if json.loads(m["outcomes"]) != ["Yes", "No"]: continue
keep(question=m["question"], rules=m["description"],
yes_token=json.loads(m["clobTokenIds"])[0],
resolved_yes=int(prices[0] == "1"), ...)
The price comes from the CLOB history of the Yes token, hourly, over the nine days before
the reference time. "24 hours before" means the last trade at or before that moment:
ref = min(end_date, closed_time) # the event, not the resolution lag
hist = get(f"https://clob.polymarket.com/prices-history?market={yes_token}"
f"&startTs={ref - 9*86400}&endTs={ref}&fidelity=60")["history"]
def last_before(hist, t):
p = None
for h in sorted(hist, key=lambda h: h["t"]):
if h["t"] <= t: p = h["p"]
else: break
return p
p_24h = last_before(hist, ref - 86400)
p_7d = last_before(hist, ref - 7 * 86400)
Markets younger than a day have no 24-hour price and are dropped. That removes most of the
daily crypto and same-day sports markets.
Asking Jev
One request per market, two questions. The state is a small JSON object, as the docs
recommend, and the Noul is phrased as a statement that is either true or false:
from typesafe_sdk import AsyncTypeSafeClient, Choice, Noul
state = {
"market_question": "Will Xi Jinping visit US by September 30?",
"resolution_rules": "This market will resolve to Yes if ...", # cut at 1,200 chars
# priced variant adds:
# "market_price_yes": 0.27,
# "market_price_note": "Price of one Yes share one day before ... the traders' implied probability",
}
questions = {
"yes": Noul(instructions=(
"The prediction market described in the state resolves Yes, meaning the event asked "
"about in market_question happens according to resolution_rules.")),
"category": Choice(
instructions="Which category does market_question belong to?",
criteria={"sports": "A sports match, game, season, tournament or athlete result",
"crypto_price": "The price of a cryptocurrency reaching, staying above or below a level",
"politics": "Elections, governments, legislation, geopolitics, wars, officials",
"economics": "Central banks, interest rates, macro data, companies, stocks, commodities",
"entertainment": "Movies, music, awards, television, celebrities, culture",
"science_tech": "Technology products, AI models, space launches, science, weather, health",
"other": "Anything that does not fit the categories above"}),
}
async with AsyncTypeSafeClient(model="jev-latest") as client:
resp = await client.system_one(state, questions)
p_yes = resp.nouls["yes"].noul # 0..1
category = resp.choices["category"].choice
The runner sends eight requests at a time and appends every answer to a JSONL file, so a
crash or a rate limit just means rerunning the same command. All 52,348 requests went
through without a single error, which is more than I can say for most APIs.
Scoring is the usual for probabilities: Brier score and log loss (lower is better), AUC
(can it rank the markets that resolved Yes above the ones that resolved No), and calibration.
Does Jev beat the price?
No, on any measure.
| Method, 9,138 markets resolved 10 Sep – 10 Oct 2026 | Brier | Log loss | AUC (95% CI) |
|---|---|---|---|
| Always predict the base rate (35% Yes) | 0.227 | 0.647 | 0.500 |
| Market price, 7 days before (5,229 markets old enough) | 0.170 | 0.498 | 0.773 |
| Market price, 24 hours before | 0.152 | 0.450 | 0.844 (0.836–0.852) |
| Jev, question and rules only | 0.220 | 0.630 | 0.615 (0.603–0.627) |
| Jev, told the 24-hour price | 0.190 | 0.556 | 0.794 (0.785–0.803) |
Jev on its own is barely better than guessing the base rate (Brier 0.220 against 0.227). The
market a full week out beats it comfortably. And when I handed Jev the price, it tracked it
(rank correlation 0.85) but dragged every answer towards its favourite region around 0.3.
Its adjustment made the squared error worse on 72.6% of markets.

The same gap holds in every category Jev assigned. It's smallest in sports (0.656 vs 0.774)
and largest where the price is almost always right, like crypto price thresholds (0.698 vs 0.971)
and politics (0.633 vs 0.956).
Jev almost never says Yes
This is the part worth looking at. Here's where the answers land:

| Median | 5th–95th percentile | Above 0.5 | |
|---|---|---|---|
| Market price, 24 h before | 0.27 | 0.00–0.94 | 23.0% of markets |
| Jev, blind | 0.32 | 0.20–0.44 | 2.4% of markets |
Jev answered between 0.2 and 0.45 for nine markets out of ten, whatever the question. A
market priced at 1 cent and a market priced at 99 cents got nearly the same number.
Within that narrow band it's reasonably calibrated: when it said 0.19 about 17% resolved
Yes, when it said 0.47 about 48% did. So it isn't lying. Given a question it can't know the
answer to, it returns something close to "a prediction market usually resolves No", which is
true (35% resolved Yes here) and useless for picking trades.

Does it add anything to the price?
Three checks, all negative.
Blend. A logistic regression on the market's logit and Jev's logit, 5-fold
cross-validated, gives Jev a weight of +0.20 and leaves the Brier score at 0.150, the same as
recalibrating the market price alone.
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import KFold
X = np.column_stack([logit(p_market), logit(p_jev)])
pred = np.zeros(len(y))
for train, test in KFold(5, shuffle=True, random_state=0).split(X):
m = LogisticRegression(C=10.0).fit(X[train], y[train])
pred[test] = m.predict_proba(X[test])[:, 1]
print(brier(pred, y), m.coef_) # 0.150, [+0.92, +0.20]
Inside a price bucket. This is the "mispricing filter" use. Take all markets priced
between 35 and 65 cents and ask whether Jev can tell which ones will resolve Yes. Its AUC
inside that bucket is 0.536. For longshots under 5 cents it's 0.480. For favourites above
85 cents it's 0.379, worse than reversing it. Once you know the price, Jev's number carries
nothing.
Paper trades. Buy a Yes share when Jev is higher than the price by more than a
threshold, a No share when it's lower. Only markets priced between 10 and 90 cents, no fees
or spread.
for d in (0.1, 0.2, 0.3, 0.4):
buy_yes = p_jev - p_market > d
buy_no = p_market - p_jev > d
pnl = np.concatenate([y[buy_yes] - p_market[buy_yes], # Yes share pays 1 or 0
(1 - y[buy_no]) - (1 - p_market[buy_no])]) # No share costs 1 - price
print(d, len(pnl), pnl.mean())
| Jev disagrees with the price by | Trades | Mean profit per share (± s.e.) | Random, same number of trades |
|---|---|---|---|
| more than 10 cents | 3,510 | +1.1¢ ± 0.8 | +0.0¢ ± 0.8 |
| more than 20 cents | 1,512 | −1.1¢ ± 1.2 | +0.0¢ ± 1.2 |
| more than 30 cents | 669 | −3.8¢ ± 1.6 | −0.1¢ ± 1.7 |
| more than 40 cents | 299 | −5.0¢ ± 2.1 | −0.2¢ ± 2.8 |
The more strongly Jev disagrees with the market, the more it loses. That's the market being
right about things Jev can't know.
What Jev gets most wrong
The largest disagreements tell the story better than the tables:
| Market | Price | Jev | Resolved |
|---|---|---|---|
| Google Maps renames Lake Ontario to "Lake America" by September 30, 2026? | 0.94 | 0.05 | Yes |
| Next US-Iran senior diplomatic meeting by September 30, 2026? | 0.05 | 0.93 | No |
| Will Xi Jinping visit US by September 24? | 1.00 | 0.14 | Yes |
| Will Kyler Murray be the Vikings' Week 1 starting QB? | 0.97 | 0.10 | Yes |
| Will "Patient Zero - Taylor Swift" be the #1 song this week? | 1.00 | 0.15 | Yes |
Every one of these is a question whose answer was in the news the day before, and not in
any text Jev was trained on. Jev's 0.05 on Lake America is a perfectly sensible prior about
the world as it was described in its training data. The market had read the news. That's
the whole gap: a prediction-market trader's edge is current information, and the question
text doesn't contain any.
Where Jev does fine
Reading. The Choice question asked which of seven categories the market belonged to, and
Jev's answer matched a keyword rule on 98.6% of the 8,105 markets the rule could label. Same
pattern as part 5: Jev reads the input accurately and
answers the question it can answer from the input. Whether an event in the future happens is
not that question.
What I couldn't verify
- One model version, jev-1.13.0, and one way of asking. I tried a single Noul phrasing.
A different phrasing moves the numbers a little; it doesn't give Jev the news. - The blind state has no context. A real bot would add news, odds from a bookmaker, or
the current BTC price. Then you're testing the context, and Jev becomes a reader of it,
which is the use I'd actually expect to work. I only tested whether Jev knows anything on
its own. - Markets under a day old are excluded, because they have no 24-hour price. Those are
mostly daily crypto and sports markets. - Paper trades ignore spread and assume you could buy at the last traded price. Real
fills would be worse, so a negative result stays negative. - Volume filter. $10k and up. Thin markets may be mispriced more often; they're also
harder to trade.
What this means for a Jev prediction-market bot
If I were building one, I'd take the pitch at its word and measure it the way it
suggests: if the scanner finds 1,000 candidates and Jev keeps 30, those 30 have to resolve
better than the 1,000. On question text alone they don't. Jev's output is a near-constant 0.3
plus noise, and acting on the noise loses money in proportion to how loudly it disagrees
with the price.
What might work is the boring version: code finds the candidates and gathers the current
information, and Jev reads that, with questions whose answers are in the text it's given.
"Does this article say the meeting was cancelled?" is a Jev question. "Will the meeting
happen?" isn't.
Everything is on GitHub so you can rerun it:
github.com/truongxxxx/jev-polymarket-test.
Five scripts, no notebook: fetch markets, fetch prices, build the dataset, call Jev, score.
The Jev part costs about $1.60 for all 26,174 markets, or a few cents for one month.
Next, part 8: I expected Jev to remember what
happened in 2025. It didn't, and that changes how you can backtest it.
Comments
No comments yet. Questions about the settings, the data or the numbers are welcome.
Log in or create a free account to comment.