Results
Measured after training.
The goal is not a better frozen model. The goal is intelligence whose capability grows after deployment.
Novel Knowledge Acquisition Rate (NKAR)
NKAR measures Primer’s ability in our experimental environment to correctly acquire rules generated after the model checkpoint is frozen.
0.852
Best demonstrated NKAR
Primer 1.3B · Internal post-freeze research benchmark
Gen I
0.000
Gen II
0.333
Gen III
0.778
Gen IV
0.852
Evaluation targets are generated after checkpoint freeze to reduce the possibility that success comes from memorized training examples.
Internal research benchmarks. External validation is ongoing. These results do not demonstrate AGI or superintelligence.
Memory
Learning isn’t useful if you forget.
PR
1.00
Retention
PI
1.00
Interference resistance
14 / 14
Retention and interference probes passed in our memory experiment.
FDR
0.00
For the tested verified-memory pathway.
In our experimental memory evaluation, Primer retained previously acquired knowledge even after intervening learning tasks.
The goal is persistent intelligence: knowledge acquired today should remain useful tomorrow.
Generalization
Harder concepts.
Latest result
43%
of newly generated structural rules correctly acquired after checkpoint freeze, on our strictest evaluation instrument. This capability was absent two research generations earlier.
Recent experiments extended post-training learning to a qualitatively harder class of concepts: structural rules rather than simple mappings. Zero false discoveries were recorded on this instrument, and previously demonstrated memory retention remained intact.
Coverage of this concept class is still incomplete. Generalization remains an active research milestone.
Internal research benchmark. External validation pending.
A previous version of this page reported 47% from an earlier evaluation. The figure was revised after an internal audit strengthened the separation between training and evaluation data. The qualitative finding was unchanged.
Composition
Can knowledge become a building block?
22 / 22
Compositional targets solved when the required prerequisite knowledge was available.
This experiment demonstrated that previously available knowledge could support successful solutions to harder related problems. Later experiments closed much of the remaining autonomy gap: the system now acquires, retains, and reuses that knowledge on its own. Generalization across broader domains remains an active research milestone.
Internal research benchmark. External validation pending.
Compounding
Does learning compound?
21 / 26
Harder targets solved through a multi-step chain of the system's own acquired knowledge. Without the chain: 0 / 26.
81% vs 59%
Success rate on later learning tasks using the system's own previously acquired knowledge, versus learning fresh — at lower cost.
The system learned a concept, used it to learn a harder one, then used that to learn a harder one still. Every step drew exclusively on knowledge from the system's own verified store, with zero false discoveries. This is the first direct demonstration of the compounding behavior Primer is built to study.
In a preregistered evaluation, knowledge the system had acquired and retained on its own made subsequent learning both more successful and cheaper than starting from scratch. The result replicated across independent seeds and checkpoints with zero false discoveries.
Internal research benchmark. External validation pending.