# Rediscovery Is Not Regression

Jev, “just a classifier,” and what happens when AI starts becoming a system

Volume 55 · Edition 2 · 2026-09-25

Canonical: https://www.thierry-gilgen-ict.ch/field-notes/rediscovery-is-not-regression

Nine days ago I wrote about a model called Jev.

The part that interested me was not the launch language around “System One Models.” It was the smaller idea underneath it: perhaps the useful unit of intelligence inside software is sometimes not an agent, a conversation, or a paragraph.

Perhaps it is a judgement.

A bounded question. A defined answer space. A probability. Then code decides what happens next.

Since then, something predictable has happened.

The honeymoon lasted approximately no time at all.

The posts appeared. Some thoughtful, some irritated, some almost offended by the attention Jev was receiving.

*It is just a classifier.*

*We had this twenty years ago.*

*Probabilistic models are not new.*

*Calibration is not new.*

*The industry forgot machine learning and has now rediscovered it.*

One version went further: this was evidence that AI was moving backwards. Perhaps even the beginning of another AI winter.

There is an uncomfortable fact inside the criticism.

A lot of it is true.

And I think the conclusion is wrong.

## Yes. We had classifiers.

Classification did not begin in September 2026.

Neither did probabilistic prediction.

Neither did confidence calibration.

In 2017, Chuan Guo and colleagues published a widely cited study of neural-network calibration, examining whether predicted probabilities correspond to observed correctness and showing that modern neural networks could be badly calibrated. Temperature scaling, one of the practical methods they evaluated, was itself a refinement of ideas that came earlier.[^2]

SetFit later showed that relatively small sentence-transformer models could be adapted efficiently for few-shot text classification without requiring a large generative model and elaborate prompts.[^11]

GLiNER demonstrated something else that matters to this discussion: a compact bidirectional model could accept arbitrary entity descriptions and perform flexible zero-shot extraction while avoiding the sequential token generation of a conversational LLM.[^12]

None of this is hidden history.

The field did not spend decades waiting for somebody to invent `if probability > 0.8`.

So if the claim were:

> Jev invented classification.

The criticism would be devastating.

But that is not the interesting claim.

The interesting question is what happens when a familiar capability crosses a new boundary of **generality, usability and economics**.

That is the layer worth looking at.

## The primitive is not the abstraction

A database did not invent storing information.

A GPU did not invent multiplication.

Unix did not invent files, processes or programs.

Containers did not invent process isolation.

Useful abstractions frequently arrive long after their underlying primitives.

What changes is the boundary.

A capability that previously required specialist work becomes accessible behind an interface other engineers can use without rebuilding the machinery underneath it.

That interface can change an industry even when almost every ingredient has ancestors.

This is why “we had this before” is such a dangerous sentence in technology.

Sometimes it is a necessary correction to marketing.

Sometimes it is evidence that we are looking at the wrong layer.

The question is not only:

> Has somebody done classification before?

Of course they have.

The more interesting questions are:

> Can one model accept a changing set of natural-language decision criteria without training a new classifier for every one?

> Can it return useful probability distributions rather than merely a plausible label?

> Can several such judgements be requested cheaply enough that software architects stop rationing them?

> Can the output become a replaceable component inside ordinary code rather than the beginning of another conversation?

Those are engineering questions.

And they are not answered by pointing at a logistic regression textbook.

## What Jev actually changed

TypeSafe introduced Jev on 15 September 2026 as an early-access model for typed probabilistic decisions. Its interface takes unstructured state and developer-defined questions, then returns bounded outputs such as yes-or-no probabilities, choices and scores.[^1]

The company calls its training method Reinforcement Learning for Calibrated Decisions, or RLCD, and says it built a new architecture and parallel sampler around this task.[^1]

Those claims deserve scrutiny.

The architecture has not been publicly disclosed in enough detail for an outsider to determine how much is genuinely new. RLCD has not been documented to the point where independent researchers can reproduce the training procedure and establish its novelty.

Sebastian Raschka made the useful distinction here. His assessment is that Jev may well be based on familiar ingredients — potentially an encoder-style architecture, calibration-oriented training, strong data and a well-designed API — while still being interesting because of how well the resulting system appears to generalise across classification tasks.[^3]

That is the position I find most convincing at this stage.

I do not know whether Jev contains a research breakthrough.

I do think it has exposed a useful system boundary.

Those are different claims.

And confusing them creates both kinds of bad analysis: hype on one side and dismissal on the other.

## Computing has been here before

General-purpose systems are wonderful because they let us do things nobody predicted when the system was built.

Then a workload becomes important.

We understand it better.

We discover where the general system wastes time, energy or money.

And specialisation returns.

The history of computing is full of this movement.

Google’s Tensor Processing Unit did not appear because multiplication had been forgotten. It appeared because machine-learning workloads had become important enough that a domain-specific architecture could execute them more efficiently than a general-purpose processor.[^4]

John Hennessy and David Patterson made domain-specific hardware a central part of their 2018 Turing Lecture on a new golden age of computer architecture. Their argument was not that general-purpose computing had been a mistake. It was that the end of easy performance scaling made hardware/software co-design and domain-specific architectures attractive again.[^5]

That is not regression.

It is workload maturity.

The same pattern exists higher in the stack.

The original Unix system was valuable partly because independent programs, files, pipes and processes could be composed through simple interfaces.[^13]

A program did not have to become the operating system.

A component could remain a component.

That sounds almost embarrassingly obvious.

Yet AI software spent several years moving in the opposite direction.

We found a model capable of writing text, summarising documents, extracting data, classifying messages, calling tools, reasoning about policies and generating code.

So naturally, we asked it to do all of those things.

Often in one prompt.

Sometimes in one enormous prompt.

The wonder of the general-purpose model encouraged a general-purpose architecture around it.

That was understandable.

It was also unlikely to be the final form.

## The model became the system

The first wave of generative AI applications often looked like this:

`input → LLM → output`

Then the prompt got longer.

Then we added retrieval.

Then tool calls.

Then validators.

Then structured outputs.

Then retries.

Then routing.

Then a second model to judge the first model.

Then a stronger model for the difficult cases.

Then code to decide which model should run.

At some point, the “model” was no longer the application.

It was one component inside an application again.

Berkeley researchers described this transition in 2024 as the shift from models to **compound AI systems**: systems where models interact with retrieval, tools, programmatic logic and other models rather than attempting to solve the entire task in one monolithic call.[^14]

That shift has continued for a simple reason.

Different parts of work have different computational shapes.

Arithmetic is not classification.

Classification is not retrieval.

Retrieval is not planning.

Planning is not permission.

Permission is not text generation.

We can certainly ask a sufficiently capable language model to approximate all of them.

That does not mean we should.

A universal model is useful because it expands the space of things software can attempt.

A specialised component becomes useful when we understand a recurring part of that space well enough to give it a sharper interface.

The two ideas are not enemies.

The general model discovers territory.

The system eventually builds roads.

## What the Bitter Lesson actually says

This is where the discussion becomes more interesting.

Rich Sutton’s *The Bitter Lesson* is one of the most important short essays in modern AI. It is also easy to turn into a slogan.

The slogan is something like:

> General always beats specialised. Scale wins. Stop engineering.

That is not Sutton’s argument.

His observation was that AI researchers repeatedly tried to encode their own understanding of a domain into systems, while more general methods that could exploit increasing computation — particularly search and learning — eventually surpassed those hand-designed approaches.[^6]

Chess is one example.

Go is another.

Speech recognition and computer vision follow similar patterns in Sutton’s account.

The bitter part was not that every system should contain one enormous model.

It was that human beings are bad at manually specifying all of the knowledge required for intelligence, while learning and search can continue improving as computation grows.

That lesson is completely compatible with specialised learned models.

A decision model trained from data is not the same thing as an expert system containing thousands of rules written by a committee.

A TPU is specialised hardware, yet it exists precisely to make large-scale learning more computationally effective.

A retrieval system is specialised, yet it may supply a general model with knowledge it would otherwise have to approximate from parameters.

A classifier can be specialised in its output space while still having acquired its capabilities through general learning.

The important distinction is not:

**general versus specialised**

but often:

**learned versus manually encoded**,  
**reusable versus brittle**,  
**scalable versus trapped by its own design**.

TypeSafe has entered this discussion directly with an essay called *The Bitterest Lesson*. Its argument is that compute matters enormously, but the objective being optimised matters first: a system can scale beautifully while becoming better at the wrong task.[^15]

That is a vendor’s argument and should be read as such.

But the underlying question is legitimate.

If the job is to choose among five permitted actions, is next-token generation the computational objective we ultimately want to optimise?

Maybe.

Maybe not.

That is an empirical question, not an article of faith.

## Then everyone built one

There is another reason I think the Jev episode is worth watching.

The response was unusually fast.

Bespoke Labs released Nimble, an open Jev-inspired model that takes text and a schema and returns typed decisions with probabilities. The project says its first version was built in one day.[^8]

Together AI released Tev1-4B-experimental on top of Qwen3.5 4B. Its published training recipe uses roughly 38,000 classification examples and reports a fine-tuning cost of about US$17.[^7]

Laya implements a non-autoregressive Jev-compatible decision interface and publishes comparative benchmarks that are useful precisely because they are not universally flattering. Its own results show areas where its fine-tuned model outperforms Jev and areas — including large choice sets and some probability-distribution metrics — where Jev remains stronger.[^9]

Kev takes the idea down another path again: small, open Jev-style decision models, including a laptop-scale prototype and larger Qwen-based variants. Its model cards explicitly warn that a confidence value is not a verified probability of correctness and that calibration should be measured on the actual deployment population.[^10]

None of these projects proves that TypeSafe has created a permanent category.

They prove something narrower.

Other builders saw the interface and immediately thought:

**I want that.**

That matters.

A technological category is not created when a company names it.

It starts becoming real when competitors, open-source developers and users can recognise the shape of the thing well enough to reproduce it, vary it and argue about it.

Perhaps “System One Model” will become the name.

Perhaps nobody will use that term a year from now.

The name is secondary.

The emerging pattern is more interesting:

**state in → bounded questions → probability distributions out → code decides what follows**

That interface now has multiple implementations.

For something allegedly so outdated, it is generating a surprising amount of new work.

## Rediscovery can still be overhyped

None of this means Jev wins.

In fact, the strongest response to the current enthusiasm is not “classifiers are old.”

It is:

> Which classifier should I use for this particular work?

If I have a stable task, a clean label set and thousands of high-quality examples, a conventional task-specific classifier may be exactly the right answer.

It may be faster.

It may be cheaper.

It may be easier to evaluate.

It may run entirely on infrastructure I control.

SetFit and years of text-classification research did not stop being useful because Jev launched.[^11]

Likewise, a general LLM may remain the better component when the answer space cannot be defined in advance, when the task requires synthesis, when reasoning across unfamiliar situations matters more than latency, or when the system genuinely needs to produce language.

And then there are the probabilities.

A decimal point creates an impression of precision that the system may not deserve.

Guo’s calibration work remains relevant because a model saying `0.92` is not enough. The operational question is whether predictions at that level are actually correct at roughly that frequency on the population where the system is being used.[^2]

Distribution shift matters.

Label design matters.

Missing alternatives matter.

Bad evidence matters.

The system can be perfectly type-safe and still choose the wrong permitted answer.

Specialisation does not remove the responsibility to test.

It makes the test easier to define.

That is an advantage only if we actually do it.

## This is not what an AI winter looks like

The phrase “AI winter” carries historical weight.

Stanford’s AI100 history describes the mid-1980s winter as a period when practical disappointment caught up with expectations, interest declined and funding dried up.[^16]

That is a very different phenomenon from developers discovering that some production workloads do not require a giant generative model.

If companies begin replacing expensive general-purpose calls with smaller learned components where the task permits it, that is not evidence that AI has failed.

It is evidence that cost matters.

Latency matters.

Reliability matters.

Architecture matters.

In other words, the technology is being subjected to the same pressures as every other technology that eventually had to leave the demonstration stage.

The first phase asks:

> What can this machine do?

The next phase asks:

> Which machine should do which part?

The second question sounds less magical.

It is also the question engineers eventually have to answer.

An industry composed of LLMs, decision models, embedding models, retrievers, vision systems, deterministic code, databases, policies and human escalation paths may look less like science fiction than one omnipotent model behind an API.

It may also work better.

## Boring is a milestone

There is a strange habit in technology discussions.

We treat generalisation as progress and specialisation as retreat.

But an invention becoming boring is often one of the clearest signs of success.

Electricity disappeared into walls.

Networking disappeared into operating systems.

Databases became infrastructure.

Virtual machines became ordinary.

Cloud computing eventually became somebody else’s invoice.

A mature technology is surrounded by components whose novelty no longer matters to the person using them.

AI will probably be no different.

There will still be frontier models.

They will become more capable.

They will open new territory.

But the systems built around them may become increasingly heterogeneous.

Some decisions will go to a frontier reasoning model.

Some will go to a four-billion-parameter classifier.

Some will go to a model running locally.

Some will go to a SQL query.

Some will go to a rules engine.

Some will go to a person because the consequence is too important or the evidence too incomplete.

Good architecture will not care which component currently has the most impressive benchmark screenshot.

It will care whether the component does its assigned work well enough, cheaply enough and predictably enough to be depended upon.

That is not a retreat from intelligence.

It is intelligence becoming infrastructure.

## Rediscovery is not regression

So yes.

Jev is a classifier.

Or, at least, classifier is a perfectly reasonable description of what it does.

Classification is old.

Calibration is old.

Specialised models are old.

None of those observations settle whether Jev is useful.

They do something better: they remove the mythology.

Then we can ask the questions that matter.

Does the interface generalise?

Are the probabilities useful?

How does it behave outside the vendor’s examples?

Where does an ordinary classifier beat it?

Where does a generative model beat it?

Can the component be replaced without rebuilding the application?

What does the whole system do when the model is wrong?

Those are much healthier questions than asking whether we have crossed some invisible historical line between “old AI” and “new AI.”

Technology rarely advances by permanently discarding everything that came before.

It advances by recombination.

A capability becomes general.

We use it everywhere.

We learn where it fits.

We specialise the expensive parts.

We standardise the useful interfaces.

Then another general-purpose breakthrough arrives and the cycle begins again.

If Jev disappears next year, that pattern will still be worth understanding.

Because the more consequential shift may have nothing to do with Jev.

It may be that the generative-model era taught developers to treat intelligence as something software can call.

And now we are beginning to ask the obvious next question:

**What kind of intelligence should each call actually require?**

That is not an AI winter.

It is architecture.

And rediscovery is not regression.

---

## Sources

Sources were accessed on 25 September 2026. Vendor documentation is used to establish product claims, interfaces and published measurements, not as independent verification of performance. Benchmarks from open-source projects are reported as project-published results unless otherwise stated.

## Footnotes

[^1]: Diogo Almeida / TypeSafe AI, “Introducing System One Models & Jev,” 15 September 2026. Launch description, System One interface, RLCD claim, architecture claim, pricing and evaluation caveats. https://typesafe.ai/blog/introducing-system-one-models-and-jev
[^2]: Chuan Guo, Geoff Pleiss, Yu Sun, Kilian Q. Weinberger, “On Calibration of Modern Neural Networks,” ICML 2017 / PMLR 70. https://proceedings.mlr.press/v70/guo17a.html
[^3]: Sebastian Raschka, “It’s Easy to Dismiss Jev as Just a Classifier,” 20 September 2026. Secondary technical analysis; explicitly notes that Jev’s exact architecture and training method are undisclosed. https://www.sebastianraschka.com/blog/2026/jev-classification-generalization.html
[^4]: Norman P. Jouppi et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” ISCA 2017. https://research.google/pubs/in-datacenter-performance-analysis-of-a-tensor-processing-unit/
[^5]: John L. Hennessy and David A. Patterson, “A New Golden Age for Computer Architecture,” Turing Lecture / ISCA 2018. https://iscaconf.org/isca2018/turing_lecture.html
[^6]: Rich Sutton, “The Bitter Lesson,” 13 March 2019. https://bitterlesson.ai/
[^7]: Hassan El Mghari / Together AI, “How to train your own Jev for $17,” 23 September 2026. Tev1-4B-experimental, Qwen3.5 4B base, published data recipe and reported fine-tuning cost. https://www.together.ai/blog/how-to-train-your-own-jev
[^8]: Bespoke Labs, “Nimble,” GitHub repository, accessed 25 September 2026. Open Jev-inspired typed-decision model; project states the initial version was built in one day. https://github.com/bespokelabsai/nimble
[^9]: NandhaKishorM, “Laya,” GitHub repository, accessed 25 September 2026. Jev-compatible non-autoregressive decision engine and project-published comparison benchmarks. https://github.com/NandhaKishorM/laya
[^10]: Jared Palmer, “Kev,” GitHub repository and model cards, accessed 25 September 2026. Open Jev-style decision models and calibration caveats. https://github.com/jaredpalmer/kev
[^11]: Lewis Tunstall et al., “Efficient Few-Shot Learning Without Prompts,” 2022. https://arxiv.org/abs/2209.11055
[^12]: Urchade Zaratiana, Nadi Tomeh, Pierre Holat, Thierry Charnois, “GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer,” NAACL 2024. https://aclanthology.org/2024.naacl-long.300/
[^13]: D. M. Ritchie and K. Thompson, “The UNIX Time-Sharing System,” Communications of the ACM / Bell System Technical Journal edition. https://people.eecs.berkeley.edu/~brewer/cs262/unix.pdf
[^14]: Matei Zaharia et al., “The Shift from Models to Compound AI Systems,” Berkeley AI Research, 18 February 2024. https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/
[^15]: TypeSafe AI, “The Bitterest Lesson,” 10 September 2026. Vendor essay arguing that task selection and data precede compute and algorithms in practical ML system design. https://typesafe.ai/blog/bitterest-lesson
[^16]: Stanford One Hundred Year Study on Artificial Intelligence, “Appendix I: A Short History of AI,” 2016 report. Historical description of the AI winter. https://ai100.stanford.edu/2016-report/appendix-i-short-history-ai

## Connected reading
- [Why AI projects fail](https://www.thierry-gilgen-ict.ch/field-notes/why-ai-projects-fail)
- [The Scan Is Not the Book](https://www.thierry-gilgen-ict.ch/field-notes/the-scan-is-not-the-book)
- [Open Is Not Sovereign](https://www.thierry-gilgen-ict.ch/field-notes/open-is-not-sovereign)
- [AI-First Is an Operating Model](https://www.thierry-gilgen-ict.ch/field-notes/ai-first-is-an-operating-model)
- [Access Is Not Authority](https://www.thierry-gilgen-ict.ch/field-notes/access-is-not-authority)
