Every model we route to can say why it is there.
An insect's elytra shield its wings. Ours shields the model estate: eight stages from open flood to approved set.
Eight stages, one closed cycle.
What makes Elytra a protocol rather than a scoreboard is that the stages are wired to each other, and that the population is always moving through them. Every model in the router has been round all eight, and can say who signed for it.
The six hour sweep
Five pipelines are swept every six hours, and every new line is admitted to the watch on the same terms. Freshness is a hard gate rather than a preference.
A model has to have moved recently to be in the window at all. A line that has stopped shipping has stopped earning attention.
rejectsWeights nobody has touched since last year.
- cadence
- every 6 h
- window
- rolling, months
The loop has no end state. What Compose ships, Evolve re-examines. What Evolve retires, Capture eventually replaces. A seat is earned continuously, because the index is a percentile against a moving population: standing still is falling.
Neural Arc · model research frameworkSeven hundred and fifty under watch. Four hundred and sixty seats.
Each cell is a model we are currently watching. Step through the gates and watch the estate thin. Nothing here falls for being bad, only for failing to prove otherwise.
Three gates sit between a model appearing and a model being routed to, and a human sits after all three. The four hundred and sixty that make it are coloured by the stage they are in now, because a seat is a place in a protocol that is still running. Select a gate to hold the field there.
A good result on thin evidence is not a good result.
Three quarters of eight is the same fraction as three quarters of two hundred, and it is not the same claim. Elytra never scores a model at what it appeared to do. It scores it at what we would be willing to defend, which is a lower number until the evidence earns it.
It is an unflattering way to run a leaderboard, and that is the reason for it. A number that goes up only when the evidence behind it goes up is one we can put in front of a client, and it is the number the rest of the protocol reads.
Match a shape, not a leaderboard position.
A model is not good or bad. It has a shape. A piece of work has a shape too, and the only question worth asking is how closely the two agree. Elytra describes both on the same axes, so the answer stops being an opinion.
Fit against this use case
2
Axes short of what the work needs
6 of 8 clear
- Instruction fidelityover by +4
- Long-context recallover by +8
- Domain reasoningshort by −22
- Open-ended rangeover by +56
- Tool and function useover by +1
- Multilingual coverageover by +23
- Latency under loadshort by −26
- Cost efficiencyover by +4
6 axes clear and 2 short. The short ones are domain reasoning and latency under load, and no amount of the axis it is enormous on buys either of them back.
Eight axes, and a model that is enormous on the wrong ones is still the wrong model.
Measurement outranks description
Where a model runs on our hardware, what it actually did is the evidence. A publisher's own account of what it is for counts for something, and never for as much, because a description is not a measurement.
The shape is authored
Each use case profile is written by the people who ship in that industry, then versioned like source code. It is a commitment rather than a guess, and every stage of the protocol reads the same one.
Capability is not the whole answer
A model also has to be the right size for the job, and it has to be deployable under a licence the client will sign. A model that scores well and cannot be run where the work is has not scored well.
One framework, six industries that disagree about it.
The mathematics is shared. The settings are not. Each industry owns a thesis about what intelligence means in that vertical, and the framework is bent to its physics and its regulation rather than to an average of everybody's.
Telecom
Networks make the densest event streams in industry. The lens triages at the edge with small models and escalates only where synthesis is the deliverable, because latency here is a correctness criterion rather than a comfort.
Manufacturing
Factory intelligence has to live on the floor: air gapped, no GPUs, and the same answer every time. No lens asks more of deployability, because the contract is with fanless industrial hardware and a thermal envelope.
Insurance
Claims are documents plus judgment, so the requirement is deliberately split between the two. A model that is mediocre at both halves scores worse in this lens than it does anywhere else in the framework.
Banking
Sovereignty is not negotiable. A model that cannot be deployed on the bank's own metal, under a licence its counsel will sign, is not a candidate at any capability. The deliverable is a defensible record rather than prose.
Consulting
The product is synthesis, so the best reasoner wins almost regardless of size. This is the lens that sees new frontier class open models first, and the one where a panel has to span several model families before it is allowed to agree.
Retail
Margin lives in the long tail of content and service. A small model shipping ten thousand product descriptions a night beats a very large one that cannot be afforded at catalogue scale. The clock is a nightly batch rather than a request.
One rule holds across all six. Sovereignty never falls to zero, even in the least regulated lens, because the promise underneath every one of them is deployment on infrastructure the client controls.
A seat is earned, and earned again.
The approved set has a fixed number of seats, which is what makes one worth having. When it is full, something retires before anything is admitted.
460
Seats in the approved set. The last of them are held back deliberately, as room for models a lens backs on conviction before the rest of the market has noticed.
1
Candidate reviewed at a time. It is slow on purpose. Every review is a test of whether the framework's answer matches a researcher's, and a run of disagreement is a bug report against the framework.
0
Models admitted commercially. There is no paid placement in the estate, because a seat that could be bought would make the other four hundred and fifty nine impossible to defend.
A model estate anyone can assemble. One that can testify is harder.
That chain is the product. It is what a client, a partner, or an auditor is actually buying when they buy the estate rather than a list of models.