Fusion Aperture
A long conversation should not need a bigger machine.
The engine inside SimpleLM, and the four things we refused.
Long context
1,000,000 tokens
Limited RAM
Half-gigabyte class
Small model
Two billion, four-bit
Training-free
No specialist
Fusion Aperture
the remaining intersection
I
Long context is sold as a rental.
Every mainstream answer to remember more is a machine you do not own. A frontier API with a meter running, or a datacenter model rented by the token. The window closes when the invoice does. That is a strange place to keep a year of someone's thinking.
The usual
A metered window, rented by the token
Aperture
On the device, with nothing metering it
II
The usual fix is to train a specialist.
Take a model, train it for length, ship a larger one. It works, and it leaves you carrying a second model that has to be fed and maintained forever. We wanted the length without the luggage.
The usual
A long-context training run and a bigger checkpoint
Aperture
No training run. A stock model, unmodified.
III
A score is not a result until you publish what it cost.
Four other systems report a figure at a million tokens. Not one of them reports the memory it took to get there. A number without its cost is half a sentence, so we print both, and we do not round the peak down to make the nicer one.
The usual
A score with no memory figure beside it
Aperture
515 MB
767 MB peak, printed rather than rounded
IV
Measure on the thing people actually hold.
A number from a datacenter tells you what a datacenter can do. It does not tell you what fits in a pocket. Every figure on this site was measured on hardware a person could be carrying right now.
The usual
Benchmarked on a rack nobody owns
Aperture
Apple silicon, two billion parameters at four-bit
Retrieval held at a million tokens. The rest is on the record.
Fusion Aperture is the long-context engine inside SimpleLM. Evaluation host: a stock Qwen3.5-2B Q4_K_M, with no long-context training. Physical memory measured on Apple silicon.