PRIMER

Results

Measured after training.

The goal is not a better frozen model. The goal is intelligence whose capability grows after deployment.

Novel Knowledge Acquisition Rate (NKAR)

NKAR measures Primer’s ability in our experimental environment to correctly acquire rules generated after the model checkpoint is frozen.

0.852

Best demonstrated NKAR

Primer 1.3B · Internal post-freeze research benchmark

  1. Gen I

    0.000

  2. Gen II

    0.333

  3. Gen III

    0.778

  4. Gen IV

    0.852

Novel knowledge acquisition improved across successive research generations.

Evaluation targets are generated after checkpoint freeze to reduce the possibility that success comes from memorized training examples.

Internal research benchmarks. External validation is ongoing. These results do not demonstrate AGI or superintelligence.

Memory

Learning isn’t useful if you forget.

PR

1.00

Retention

PI

1.00

Interference resistance

14 / 14

Retention and interference probes passed in our memory experiment.

FDR

0.00

For the tested verified-memory pathway.

In our experimental memory evaluation, Primer retained previously acquired knowledge even after intervening learning tasks.

The goal is persistent intelligence: knowledge acquired today should remain useful tomorrow.

Generalization

Harder concepts.

Latest result

43%

of newly generated structural rules correctly acquired after checkpoint freeze, on our strictest evaluation instrument. This capability was absent two research generations earlier.

Recent experiments extended post-training learning to a qualitatively harder class of concepts: structural rules rather than simple mappings. Zero false discoveries were recorded on this instrument, and previously demonstrated memory retention remained intact.

Coverage of this concept class is still incomplete. Generalization remains an active research milestone.

Internal research benchmark. External validation pending.

A previous version of this page reported 47% from an earlier evaluation. The figure was revised after an internal audit strengthened the separation between training and evaluation data. The qualitative finding was unchanged.

Composition

Can knowledge become a building block?

22 / 22

Compositional targets solved when the required prerequisite knowledge was available.

This experiment demonstrated that previously available knowledge could support successful solutions to harder related problems. Later experiments closed much of the remaining autonomy gap: the system now acquires, retains, and reuses that knowledge on its own. Generalization across broader domains remains an active research milestone.

Internal research benchmark. External validation pending.

Compounding

Does learning compound?

21 / 26

Harder targets solved through a multi-step chain of the system's own acquired knowledge. Without the chain: 0 / 26.

81% vs 59%

Success rate on later learning tasks using the system's own previously acquired knowledge, versus learning fresh — at lower cost.

The system learned a concept, used it to learn a harder one, then used that to learn a harder one still. Every step drew exclusively on knowledge from the system's own verified store, with zero false discoveries. This is the first direct demonstration of the compounding behavior Primer is built to study.

In a preregistered evaluation, knowledge the system had acquired and retained on its own made subsequent learning both more successful and cheaper than starting from scratch. The result replicated across independent seeds and checkpoints with zero false discoveries.

Internal research benchmark. External validation pending.