2026 · RECOMMENDER SYSTEMS
VERDANT
A falsifiable recommendation engine for sparse, subjective preferences.
PAUSED AFTER v0.8 CONFIRMATION
"A system that cannot reject itself isn't much of an experiment."
VERDANT started with an awkward problem in Fallow, our hobby-discovery project. Most recommenders have behaviour to work with. You watched this, bought that, saved five songs, abandoned three carts. Hobby discovery often has none of it. The useful recommendation may be something you have never tried, searched for, or had a name for.
We wanted to know how much a system could infer from a few uncertain answers without pretending those answers were a complete picture of a person.
That question grew into VERDANT: a Bayesian recommendation engine that keeps preferences, constraints, motivation, and friction separate; maintains uncertainty over a user's latent state; chooses assessment questions adaptively; and scores activities using a structured utility model.
By version 0.8, VERDANT had also acquired something more important than another scoring function: a way to prove us wrong.
VERDANT / 0.8
- DEVELOPMENT_08
- COMPLETE
- SCREENS
- 05 / 11
- VALIDITY FAILURES
- 00
- AUDIT
- WITHHELD
- STATE
- PAUSED
- P
- F
- F
- F
- P
- F
- F
- P
- F
- P
- P
01 / METHOD
The cold start is the point
Fallow asks people about the kind of activity they may enjoy before it has years of clicks to study. That makes the usual recommendation shortcut less useful. A person can tell us that they like structure, or want something social, or hate long commutes. They cannot reliably hand us a ranked list of future identities.
VERDANT models that uncertainty instead of sanding it away.
Its production representation uses 26 latent coordinates:
z = [p, r, m] ∈ R²⁶q(z) = N(μ, Σ)Ucore = F · ∛(σ(E) · σ(A) · σ(I))ADAPTIVE QUESTIONS
where p carries preference state, r carries friction, and m carries motivation. The inference layer maintains a Gaussian approximation,
q(z) = N(μ, Σ)
and updates it from pairwise and structured assessment evidence with MAP estimation and a Laplace approximation.
Activities carry uncertain annotations too. The system does not treat a hobby label as a fact carved into stone. It separates estimated fit from barriers, keeps hard feasibility distinct from soft friction, and allows context to change without silently rewriting long-term preference.
For v0.8 we changed the core recommendation score from a rank-based construction to a cardinal one:
Ucore = F · ∛(σ(E) · σ(A) · σ(I))
The three terms represent the engine's modeled enjoyment, adoption, and impact channels; F applies feasibility. The geometric form makes a serious weakness in one channel matter instead of letting strength elsewhere hide it.
02 / SYNTHETIC LAB
We built a simulator because we did not have the right human data
There is an easy way to make a recommendation model look impressive: invent a benchmark that likes it.
We wanted the simulator to do the opposite.
VERDANT's synthetic lab contains 600 activities, 12 activity families, seven simulated annotators, a 260-question bank, and eleven worlds that perturb different assumptions. Some worlds add noise. Some shift preferences. Some corrupt answers. Others attack the assumptions behind the activity model.
Synthetic users are not evidence that VERDANT understands real people. We use them for narrower questions: does the algorithm behave the way we specified, does a component improve an objective when the truth is known, and does that improvement survive controlled forms of misspecification?
We froze that distinction into the project rules. Synthetic results could tune software and test scientific implementation. They could not become claims about human psychology.
03 / PREREGISTRATION
The test was frozen before we saw the answer
The first seven gates built the simulator, baselines, inference machinery, recommendation policy, metrics, deterministic random-number scheme, and validity rules. Version 0.8 then made a small set of explicit scientific bets.
We froze the code at:
C_DEV_08 = 59ebb9bf5240ff5934c1bd9bfe157915af3f3c02
Then we opened a fresh development namespace that the implementation work had never consumed.
The run executed:
- 29,696 scientific tasks
- 512 deterministic stopped-policy replays
- 11 preregistered confirmation screens
- 0 execution failures
- 0 validity failures
- 0 nonfinite scientific values
- 0 hard-feasibility violations
The scientific runner took about 7 hours 46 minutes on four local workers.
A separate AUDIT namespace was waiting behind the confirmation screen. We were allowed to open it only if all eleven conditions passed.
- DESIGN EVIDENCE
- FREEZE v0.8
- DEVELOPMENT_0829,696 TASKS + 512 REPLAYS
- 11 FROZEN SCREENS
- AUDITWITHHELD
04 / CONFIRMATION
Five passed. Six did not.
The result arrived without an implementation excuse attached to it. SCREEN-11 passed cleanly, so the failed confirmation was not a corrupted run hiding behind scientific language.
05 PASS06 FAIL
- 01CARDINAL COREPASS
- 02RAW IG AURFAIL
- 03RAW IG Q_EQUIVFAIL
- 04F3 VS F0FAIL
- 05GATE CPASS
- 06MISSPECIFICATION / GATE EFAIL
- 07SIX-ANSWER ONE-SHOTFAIL
- 08DRIFT RECOVERYPASS
- 09D5 DIVERSITYFAIL
- 10SERENDIPITYPASS
- 11INTEGRITYPASS
The cardinal score survived
Our clearest positive result came from M1.
On W0, lower regret is better:
- W0CARDINAL BETTER
- W1CARDINAL BETTER
- W2CARDINAL BETTER
- W3CARDINAL BETTER
- W4CARDINAL BETTER
- W5CARDINAL BETTER
- W6CARDINAL BETTER
- W7CARDINAL BETTER
- W8CARDINAL BETTER
- W9CARDINAL BETTER
- W10CARDINAL BETTER
The paired mean difference was -0.0150536215, about an 11% reduction from the historical rank formulation.
More important, the direction held in every simulated world from W0 through W10. Cardinal scoring beat rank scoring 11 out of 11 times.
We had changed one of VERDANT's most basic assumptions, then watched the new formulation survive every misspecification world in the confirmation set. That part earned its complexity.
Our preferred question selector lost
We expected raw information gain to outperform the adjusted selector we had used before. Fresh evidence disagreed.
RAW_IG
- AUR
- 0.1359898122
- Q_EQUIV
- 7
ADJUSTED_IG_08
- AUR
- 0.1290738506
- Q_EQUIV
- 6
Adjusted IG reached the fixed F2 reference one question sooner and produced lower area-under-regret.
SCREEN-2 and SCREEN-3 failed for the same reason. We preferred the wrong selector.
Nothing broke. The experiment changed our mind.
Five questions was too eager
The frozen stopping search selected:
τregret = 0.158403035926751
τdecision = 0.02187763557997514
The median user stopped after 5 questions.
That sounds efficient until you look at the recommendation regret:
05QUESTIONS
SCREEN-4 required stopped F3 to beat a limit of 0.1200854969. It missed by a lot. The stopping selector did what we told it to do; the frozen objective allowed a policy that traded away too much recommendation quality for fewer questions.
The miss was not confined to W0. Across the eight Gate-E misspecification worlds, F3 beat F0 in 0 of 8 worlds. Median F3 regret was 0.1606971848, versus 0.1528195242 for F2.
The five-role slate was expensive
VERDANT's slate policy tried to do more than return the five highest-utility activities. It assigned semantic jobs to recommendations, including an Easy Win, Serendipity, and Exploration.
Fresh ablations made the cost visible.
- D5 USERS
- 512 / 512
- REDUNDANCY
- 0.000
- UTILITY RETENTION
- 94.47%
- REQUIRED
- 97%
Every tested replacement improved regret.
The D5 screen adds another angle. All 512 W0 users had at least five primary-eligible activities, and mean redundancy was 0.0. The slate was diverse. It still retained only about 94.47% of the pure-top-five comparator's true utility, below the frozen 97% requirement.
We had succeeded at making the slate structurally interesting. We had not shown that the structure was worth what it cost.
05 / SURVIVORS
Two unusual ideas did survive
The drift test passed for a reason the old metric would have missed.
After a simulated preference shock:
- B_PRE0.1596346085
- L0 / SHOCK0.1680195819
- L10 / RECOVERY0.1093558435
- DAMAGE RECOVERY
- 6.996×
- LEGACY RATIO
- 0.349
- LEGACY
- REPORTING ONLY
410/ 512
- OBSERVED
- 80.078%
- THRESHOLD
- 80%
The shock caused 0.0083849734 of measured damage. Ten new answers recovered far more than that initial loss, producing an uncapped damage-recovery value of 6.9963.
The older reporting ratio was only 0.3491. Version 0.8 had frozen the corrected damage-based definition before the run, and that definition passed.
Serendipity passed too, by one of the narrowest margins available at this population size: 410 of 512 eligible users succeeded, or 0.80078125, against a threshold of 0.80.
We would not move either threshold if the number landed on the other side. Passing only means the frozen rule passed.
06 / CLOSED GATE
So we did not run AUDIT
Version 0.8 failed six of eleven confirmation screens.
Our frozen rule said AUDIT could run only if every screen passed and the integrity checks were clean. The integrity checks were clean. The scientific conditions were not.
We left AUDIT untouched.
That choice matters more to us than finding a flattering way to summarize the scorecard. The whole point of separating design evidence, confirmation evidence, and a held-out audit was to stop ourselves from treating each disappointing result as permission for one more tweak.
The DEVELOPMENT_08 realization is now design evidence. We will not rerun it as a fresh confirmation attempt, and we will not retune version 0.8 against it.
07 / CURRENT STATE
Current state
VERDANT is paused after v0.8 confirmation.
We are keeping the parts that earned our confidence, especially cardinal scoring and the discipline around uncertainty, drift, validity, and reproducibility. The confirmation run also gave us clear reasons to question raw information gain, aggressive stopping, and the current five-role composition.
Fallow can use those lessons without pretending VERDANT 0.8 graduated.
We expect to return to the research. A future version would start as a new hypothesis with a new preregistration and a fresh development namespace. It would not be a patched interpretation of this result.
For now, the experiment is finished.
PAUSED AFTER v0.8 CONFIRMATION
- C_DEV_08
- 59ebb9bf…
- SUMMARY
- 503c3d6e…
- AUDIT
- UNTOUCHED