Two weeks ago I wrote about the hallucination question. Someone from a large GSI asked which models we use. I said we have the context, so we use a dumber model.
That same week a company shipped a model that can’t write a sentence at all.
I haven’t touched it. I’ve been reading what people are doing with it, and the thing I notice is that they’re solving the problem the same way we did. Connect the data. Give the model a small job. Ask it how sure it is. Nobody is asking which model is safe.
The model is called Jev. You hand it a block of context and a list of typed questions. It answers the questions and gives a number for how sure it is. That is the whole product. No prompt. No paragraph. No box.
It refuses to generate text. Last week I wrote that a chat box is the raw shape of a language model with nothing designed in front of it. This is the other end. Someone designed the talking out.
The person who built it worked on the training method most chat models use. He has said that method rewards sounding right. His fix was to train a model that knows how sure it should be instead. I can’t check that claim. Nobody can yet. The company hasn’t published the architecture or a paper. Every number is theirs.
I’m not writing about the numbers.
I’m writing about where two roads ended up.
We got to a small model by watching cost. We had to connect the data first because the cheap model couldn’t cover for a vague job. By the time the model got involved, the job was small. Read this. Answer from it.
They got to a small model from another angle. They looked at how models are trained and decided the problem was the model’s confidence, not its size. So they built one that only decides and tells you how much to trust the decision.
Same place. Different road.
That is what I keep coming back to. The work I described last month, finding what belongs together, cutting what the model doesn’t need to read, handing it the smallest job left, is turning into a product category. Someone looked at that work and decided it deserved its own model.
The problem was never the model.
Last time, I said I don’t have a test proving the small model is enough. I have a design that leaves it very little to get wrong. That hasn’t changed. What might change is what the small job gets handed to.
It’s new. The concept is there. We want to test it before we switch anything.
The GSI person wanted a model name. I gave them a list of work. Other people made the same list.