Skip to content
FI

2026 · RECOMMENDER SYSTEMS

VERDANT

A falsifiable recommendation engine for sparse, subjective preferences.

PAUSED AFTER v0.8 CONFIRMATION

"A system that cannot reject itself isn't much of an experiment."

VERDANT started with an awkward problem in Fallow, our hobby-discovery project. Most recommenders have behaviour to work with. You watched this, bought that, saved five songs, abandoned three carts. Hobby discovery often has none of it. The useful recommendation may be something you have never tried, searched for, or had a name for.

We wanted to know how much a system could infer from a few uncertain answers without pretending those answers were a complete picture of a person.

That question grew into VERDANT: a Bayesian recommendation engine that keeps preferences, constraints, motivation, and friction separate; maintains uncertainty over a user's latent state; chooses assessment questions adaptively; and scores activities using a structured utility model.

By version 0.8, VERDANT had also acquired something more important than another scoring function: a way to prove us wrong.

VERDANT / 0.8

DEVELOPMENT_08
COMPLETE
SCREENS
05 / 11
VALIDITY FAILURES
00
AUDIT
WITHHELD
STATE
PAUSED

01 / METHOD

The cold start is the point

Fallow asks people about the kind of activity they may enjoy before it has years of clicks to study. That makes the usual recommendation shortcut less useful. A person can tell us that they like structure, or want something social, or hate long commutes. They cannot reliably hand us a ranked list of future identities.

VERDANT models that uncertainty instead of sanding it away.

Its production representation uses 26 latent coordinates:

z = [p, r, m] ∈ R²⁶
26D LATENT STATE
q(z) = N(μ, Σ)
MAP + LAPLACE
Ucore = F · ∛(σ(E) · σ(A) · σ(I))
STRUCTURED UTILITY

ADAPTIVE QUESTIONS

where p carries preference state, r carries friction, and m carries motivation. The inference layer maintains a Gaussian approximation,

q(z) = N(μ, Σ)

and updates it from pairwise and structured assessment evidence with MAP estimation and a Laplace approximation.

Activities carry uncertain annotations too. The system does not treat a hobby label as a fact carved into stone. It separates estimated fit from barriers, keeps hard feasibility distinct from soft friction, and allows context to change without silently rewriting long-term preference.

For v0.8 we changed the core recommendation score from a rank-based construction to a cardinal one:

Ucore = F · ∛(σ(E) · σ(A) · σ(I))

The three terms represent the engine's modeled enjoyment, adoption, and impact channels; F applies feasibility. The geometric form makes a serious weakness in one channel matter instead of letting strength elsewhere hide it.

02 / SYNTHETIC LAB

We built a simulator because we did not have the right human data

There is an easy way to make a recommendation model look impressive: invent a benchmark that likes it.

We wanted the simulator to do the opposite.

VERDANT's synthetic lab contains 600 activities, 12 activity families, seven simulated annotators, a 260-question bank, and eleven worlds that perturb different assumptions. Some worlds add noise. Some shift preferences. Some corrupt answers. Others attack the assumptions behind the activity model.

Synthetic users are not evidence that VERDANT understands real people. We use them for narrower questions: does the algorithm behave the way we specified, does a component improve an objective when the truth is known, and does that improvement survive controlled forms of misspecification?

We froze that distinction into the project rules. Synthetic results could tune software and test scientific implementation. They could not become claims about human psychology.

03 / PREREGISTRATION

The test was frozen before we saw the answer

The first seven gates built the simulator, baselines, inference machinery, recommendation policy, metrics, deterministic random-number scheme, and validity rules. Version 0.8 then made a small set of explicit scientific bets.

We froze the code at:

C_DEV_08 = 59ebb9bf5240ff5934c1bd9bfe157915af3f3c02

Then we opened a fresh development namespace that the implementation work had never consumed.

The run executed:

  • 29,696 scientific tasks
  • 512 deterministic stopped-policy replays
  • 11 preregistered confirmation screens
  • 0 execution failures
  • 0 validity failures
  • 0 nonfinite scientific values
  • 0 hard-feasibility violations

The scientific runner took about 7 hours 46 minutes on four local workers.

A separate AUDIT namespace was waiting behind the confirmation screen. We were allowed to open it only if all eleven conditions passed.

EXPERIMENT FLOW / FROZEN ORDER
  1. DESIGN EVIDENCE
  2. FREEZE v0.8
  3. DEVELOPMENT_0829,696 TASKS + 512 REPLAYS
  4. 11 FROZEN SCREENS
  5. AUDITWITHHELD

04 / CONFIRMATION

Five passed. Six did not.

The result arrived without an implementation excuse attached to it. SCREEN-11 passed cleanly, so the failed confirmation was not a corrupted run hiding behind scientific language.

05 PASS06 FAIL

  1. 01CARDINAL COREPASS
  2. 02RAW IG AURFAIL
  3. 03RAW IG Q_EQUIVFAIL
  4. 04F3 VS F0FAIL
  5. 05GATE CPASS
  6. 06MISSPECIFICATION / GATE EFAIL
  7. 07SIX-ANSWER ONE-SHOTFAIL
  8. 08DRIFT RECOVERYPASS
  9. 09D5 DIVERSITYFAIL
  10. 10SERENDIPITYPASS
  11. 11INTEGRITYPASS

The cardinal score survived

Our clearest positive result came from M1.

On W0, lower regret is better:

M1 / MEAN R@5LOWER IS BETTER
CARDINAL CORE0.1222066422
HISTORICAL RANK CORE0.1372602637
  1. W0CARDINAL BETTER
  2. W1CARDINAL BETTER
  3. W2CARDINAL BETTER
  4. W3CARDINAL BETTER
  5. W4CARDINAL BETTER
  6. W5CARDINAL BETTER
  7. W6CARDINAL BETTER
  8. W7CARDINAL BETTER
  9. W8CARDINAL BETTER
  10. W9CARDINAL BETTER
  11. W10CARDINAL BETTER

The paired mean difference was -0.0150536215, about an 11% reduction from the historical rank formulation.

More important, the direction held in every simulated world from W0 through W10. Cardinal scoring beat rank scoring 11 out of 11 times.

We had changed one of VERDANT's most basic assumptions, then watched the new formulation survive every misspecification world in the confirmation set. That part earned its complexity.

Our preferred question selector lost

We expected raw information gain to outperform the adjusted selector we had used before. Fresh evidence disagreed.

M2 / SELECTORLOWER IS BETTER

RAW_IG

AUR
0.1359898122
Q_EQUIV
7

ADJUSTED_IG_08

AUR
0.1290738506
Q_EQUIV
6

Adjusted IG reached the fixed F2 reference one question sooner and produced lower area-under-regret.

SCREEN-2 and SCREEN-3 failed for the same reason. We preferred the wrong selector.

Nothing broke. The experiment changed our mind.

Five questions was too eager

The frozen stopping search selected:

τregret = 0.158403035926751
τdecision = 0.02187763557997514

The median user stopped after 5 questions.

That sounds efficient until you look at the recommendation regret:

STOPPING / W0 POLICYLOWER IS BETTER

05QUESTIONS

F0 ORACLE BASELINE0.1264057862
F2 @ 13 QUESTIONS0.1506734973
STOPPED F30.1596346085

SCREEN-4 required stopped F3 to beat a limit of 0.1200854969. It missed by a lot. The stopping selector did what we told it to do; the frozen objective allowed a policy that traded away too much recommendation quality for fewer questions.

The miss was not confined to W0. Across the eight Gate-E misspecification worlds, F3 beat F0 in 0 of 8 worlds. Median F3 regret was 0.1606971848, versus 0.1528195242 for F2.

The five-role slate was expensive

VERDANT's slate policy tried to do more than return the five highest-utility activities. It assigned semantic jobs to recommendations, including an Easy Win, Serendipity, and Exploration.

Fresh ablations made the cost visible.

ROLE ABLATIONS / MEAN R@5LOWER IS BETTER
FULL FIVE-ROLE SLATE0.1596346085
REPLACE EASY WIN0.1512216445
PURE TOP FIVE0.1505070930
REPLACE EXPLORATION0.1522520257
D5 USERS
512 / 512
REDUNDANCY
0.000
UTILITY RETENTION
94.47%
REQUIRED
97%

Every tested replacement improved regret.

The D5 screen adds another angle. All 512 W0 users had at least five primary-eligible activities, and mean redundancy was 0.0. The slate was diverse. It still retained only about 94.47% of the pure-top-five comparator's true utility, below the frozen 97% requirement.

We had succeeded at making the slate structurally interesting. We had not shown that the structure was worth what it cost.

05 / SURVIVORS

Two unusual ideas did survive

The drift test passed for a reason the old metric would have missed.

After a simulated preference shock:

DRIFT / W0LOWER IS BETTER
  1. B_PRE0.1596346085
  2. L0 / SHOCK0.1680195819
  3. L10 / RECOVERY0.1093558435
DAMAGE RECOVERY
6.996×
LEGACY RATIO
0.349
LEGACY
REPORTING ONLY
SERENDIPITY / NARROW PASS

410/ 512

OBSERVED
80.078%
THRESHOLD
80%

The shock caused 0.0083849734 of measured damage. Ten new answers recovered far more than that initial loss, producing an uncapped damage-recovery value of 6.9963.

The older reporting ratio was only 0.3491. Version 0.8 had frozen the corrected damage-based definition before the run, and that definition passed.

Serendipity passed too, by one of the narrowest margins available at this population size: 410 of 512 eligible users succeeded, or 0.80078125, against a threshold of 0.80.

We would not move either threshold if the number landed on the other side. Passing only means the frozen rule passed.

06 / CLOSED GATE

So we did not run AUDIT

Version 0.8 failed six of eleven confirmation screens.

Our frozen rule said AUDIT could run only if every screen passed and the integrity checks were clean. The integrity checks were clean. The scientific conditions were not.

We left AUDIT untouched.

That choice matters more to us than finding a flattering way to summarize the scorecard. The whole point of separating design evidence, confirmation evidence, and a held-out audit was to stop ourselves from treating each disappointing result as permission for one more tweak.

The DEVELOPMENT_08 realization is now design evidence. We will not rerun it as a fresh confirmation attempt, and we will not retune version 0.8 against it.

07 / CURRENT STATE

Current state

VERDANT is paused after v0.8 confirmation.

We are keeping the parts that earned our confidence, especially cardinal scoring and the discipline around uncertainty, drift, validity, and reproducibility. The confirmation run also gave us clear reasons to question raw information gain, aggressive stopping, and the current five-role composition.

Fallow can use those lessons without pretending VERDANT 0.8 graduated.

We expect to return to the research. A future version would start as a new hypothesis with a new preregistration and a fresh development namespace. It would not be a patched interpretation of this result.

For now, the experiment is finished.

PAUSED AFTER v0.8 CONFIRMATION

C_DEV_08
59ebb9bf…
SUMMARY
503c3d6e…
AUDIT
UNTOUCHED