Faragiri

Learn AI

A model returns the plausible answer, not the true one. The eight topics here are eight distances between that answer and somebody acting on it.

7Sections
10Min read
1Depth

The division

How far a wrong answer gets before a person reads it

Four distances, and the check that works at one of them is the wrong check at the next.

  1. 04Nobody reads any single outputMachine learning and data. A total comes out, and one wrong answer inside it is invisible and still moves the total.What it was fitted to
  2. 03A stranger reads it and you never doBuilding. Attention does not scale to strangers, so the check has to be written down and run again after every change.A written case set
  3. 02A loop acted, and you read the reportAgents. The wrong answer already happened, and the report of it was written by the thing being asked.Set before it runs
  4. 01You read every wordPrompting, and generating pictures, video and audio. A wrong answer costs the minute it takes to catch it.Your attention

It says nothing about how likely a wrong answer is at each rung, which does not change much between them. What changes is who is standing there when it arrives.

A model does not return the true answer. It returns the answer that looks most like the answers it was trained on, and on most questions those are the same thing, which is what makes the questions where they are not so expensive.

Eight topics sit under this pillar, and what separates them is not which kind of media is involved or which company made the model. It is how far a wrong answer gets before a person reads it.

The division is by distance, not by modality

Almost everything published about this subject is filed one of two ways. By modality, with text in one place and images in another. Or by product, a page for each assistant, rewritten whenever a version number moves. The first writes the same advice about checking an output four times over, and the second has a shelf life of about a quarter.

Distance is the filing that survives, because the distance decides the check. Attention is enough while you are reading every word, and stops being enough the moment something acts on your behalf. A written set of test cases is overkill for one person drafting one email, and the only option once strangers are reading the output.

Two of the eight are not on that ladder, and the split is better for admitting it. Enough of the mechanism to predict the behaviour runs underneath all four rungs, because it answers why the plausible answer is the default one. Where it actually changes the job sits beside them, its subject being the person rather than the output. Of the eight pillars under Knowledge has a shape, this one dates fastest, so every dated claim is kept in one section.

Prompting first, agents last, and the tempting order is the wrong one

The order that works starts on the bottom rung and does not skip. Prompting first, and inside it, judging what came back, because everything above it multiplies whatever judgement you brought with you. Then the mechanism, which read cold is a glossary and read after some odd behaviour is an explanation. Then agents, then building, with the evaluation set before the interesting parts of it.

The wrong order is the exciting one, and it is specific: an agent before a judgement. An agent is a machine for producing outputs nobody reads one at a time, so building one before you can tell a good output from a bad one automates a judgement you do not yet have. Nothing crashes. A schedule runs, the wrong answers pile up where nobody is looking, and the discovery gets made by a customer rather than by you.

Nine jobs

The biggest ticket at Kestrel Cycles is the worst hour of work in it

Ticket price and hourly rate put the same five jobs in almost opposite orders, and only one of the two orders is visible without dividing.

Total quoted over total hours, from the nine jobs in the prompt above. The shop is invented; the division is not.

  1. Brake bleedOne job, half an hour£70.00
  2. PunctureTwo jobs, and the cheapest ticket in the shop£48.00
  3. Gear serviceTwo jobs, and the second took twice the first£40.00
  4. Wheel buildTwo jobs, six and a half hours between them£36.92
  5. Strip and rebuildTwo jobs, and the one Priya is proudest of£31.52

Nine rows is too few to price a shop on. It is enough to show that the ranking by ticket and the ranking by hour are not the same ranking, which is the part a glance gets wrong.

A prompt taken apart, on nine repair jobs

Priya Raval runs Kestrel Cycles, a two person bike shop in Bristol, and wants to know which job to stop quoting at its price. Last month is in a notebook: nine jobs, with what she charged and how long each took. Here is the prompt, whole.

  • You are pricing repair work for a two person bike shop.
  • Return one row per job type, with these columns: job type, jobs done, total quoted, total hours, price per hour.
  • Show each division on its own line before the table, so the arithmetic can be checked.
  • Then name the job type to stop quoting at this price, and the single number that decided it.
  • Last month at Kestrel Cycles: gear service, 45 pounds, 0.75 hours; gear service, 45 pounds, 1.5 hours; wheel build, 120 pounds, 2.5 hours; wheel build, 120 pounds, 4 hours; puncture, 12 pounds, 0.25 hours; puncture, 12 pounds, 0.25 hours; brake bleed, 35 pounds, 0.5 hours; strip and rebuild, 260 pounds, 9 hours; strip and rebuild, 260 pounds, 7.5 hours.
  • If a row is ambiguous, say so instead of choosing for me.

The first line only sets vocabulary, and it is the line prompt templates spend their whole budget on. The second earns its place: give the model the output format before the task. A model commits to its output one piece at a time, so the shape it starts inside is the shape it finishes in.

The third line is the point of the whole prompt. Ask for the division and the answer is falsifiable in twenty seconds with a phone calculator; leave it out and the table arrives with no way to tell which cell was computed and which was produced. The fourth line points the conclusion at a number rather than a paragraph, and the last removes the pressure to be clean where nine rows do not support it.

The arithmetic, done properly, is a total over a total. Brake bleeds: 35 pounds over half an hour, £70.00 an hour. Punctures: 24 pounds over half an hour, £48.00 an hour. Gear services: 90 pounds over 2.25 hours, £40.00 an hour. Wheel builds: 240 pounds over 6.5 hours, £36.92 an hour. Strip and rebuilds: 520 pounds over 16.5 hours, £31.52 an hour. So the biggest ticket on the list is the worst hourly rate, and it is the job Priya is proudest of.

Run the same request without the third line and a fluent paragraph comes back recommending she stop doing punctures. A puncture is twelve pounds, the smallest number on the page, and that answer is available to anyone who has divided nothing. Punctures are the second best hour in the shop.

Where the question is a decision rather than a calculation, one line changes: ask for three options and a recommendation, not an answer. Three options can be compared with the rows; a single answer can only be agreed with. That is the same skill as briefing a competent stranger, which is why being understood, in every form the job takes is a nearer neighbour than most of the software is.

The stall is believing a confident answer

People stop making progress here at the point where the answers are good enough that checking them feels like spending the time the model just saved. It does not feel like a stall. Output goes up, the week gets easier, and the first real cost arrives weeks later, inside something already sent.

The skill that gets past it is narrower than scepticism, which is what most people substitute for it. Knowing that a model can be wrong changes nothing about a Tuesday. Knowing which kinds of question it is reliably wrong about changes every Tuesday, and there are about five of them.

Anything carrying a precise number: arithmetic over rows you pasted in, a count, a date subtracted from another date. Anything very recent: a price, a version, whether a feature still exists this month. Anything about the model itself, including whether it did the thing it just said it did. Anything whose honest answer is that there is no single one, because the common case is what it was fitted to and the exception is usually what you were asking about.

The fifth costs the most. A model is most wrong where the plausible answer and the true one look alike, and nothing in the wording marks the difference: a citation that does not exist has the shape of one that does, and a function name a library never had reads exactly like one it has. It follows that confidence in the wording carries no information about correctness, because the fluent answer and the correct answer come off the same machinery.

The stall

What can be checked from the answer alone, and what cannot

Every row here arrives in the same tone, which is the reason the list has to exist at all.

What an answer tells you about itself

  1. YesThat the terms are used correctlyThe surface is what the training rewarded, and it is reliable.Yes
  2. NoThat the arithmetic was done rather than producedThe digits come off the same machinery as the sentences.No
  3. NoThat a source it cites existsA wrong citation has the shape of a right one.No
  4. NoThat it is current as of this monthThe training stopped before this month did.No
  5. Not checkedWhether this particular answer is one of the wrong onesNothing in the wording separates them, and nobody looked.Not checked

The third state is the honest one. Sorted into yes or no, that last row would be claiming somebody had checked.

What to open in 2026, and what each one replaced

A frontier assistant, meaning Claude, ChatGPT or Gemini, is what most readers already have open. As of 2026 all three replaced the search result you had to read in order to find the answer inside it, and all three have a free tier. The mistake is judging the family by that tier, a smaller model on a shorter leash than the one the company sells.

Ollama is how a model runs on your own machine, as of 2026. It replaced compiling llama.cpp by hand and hunting for weights, it is free and open source, and the cost is disk and memory. Running something small on a laptop, finding it weak and concluding models are weak is the first mistake; reading open weights as open permission is the second, because several licences forbid commercial use outright.

Claude Code, and the coding agents beside it, replaced the inline completion that finished the line you were typing, which is what tutorials written before 2024 still show. There is no free path worth the name; it runs against a paid plan or against API credits. The mistake is pointing one at a directory that is not under version control, which turns every edit into a change nobody can undo.

Building

The part of a working system that is the model is the part you can see

Everything above the line is what a demo shows and what a tutorial covers. Everything below it is what decides whether the thing survives a month.

  • The call to the modelA few lines, and the part every tutorial is about
  • The promptRewritten more often than anything else here, and cheap to change
  • The answer on the screenWhat gets demonstrated, and what gets funded
What the demo shows
  • The documents, cleanedRetrieval is a data problem wearing a model's clothes
  • The evaluation setThe only way to know a change helped rather than felt better
  • What it is allowed to touchDecided before it runs, because afterwards there is only a report
  • The log of what it actually saidNobody keeps this until the first time somebody asks
  • The licence on the weightsOpen to download is not open to use for anything

It leaves out the ratio, which nobody has measured. The drawing is a claim about where attention goes, not about lines of code.

The signal is asking twice, not asking better

At one month a person asks a question and reads the answer. At three months the same person asks the question in a form that makes a wrong answer visible, which is a different sentence rather than a better attitude.

The checkable version takes four minutes. Ask the same question twice in two fresh sessions and compare. Where the answers agree, the model is repeating something dense in its training. Where they differ, on a number, a name or a date, that is the part it was producing rather than recalling. Run it on Priya's nine rows and the table comes back twice the same; run it on what a library's function did three versions ago and it will not.

The second signal is retrospective. Take an answer you acted on this week and name the one sentence you should have checked first. If they all look equally solid, the reading is still at the level of prose, where it was at one month.

The eight topics, and which rung each is on

Getting a useful answer instead of a plausible one is where to start, whatever brought you here, because it is the only topic whose return does not depend on any of the others. The mechanism goes second.

A model given tools and allowed to act is the second rung, where the risk changes shape rather than size. An agent holding a shell, a browser and a set of credentials is a new attack surface as much as a new capability, and what it may touch has to be settled before it runs rather than read about afterwards. That is why defending systems, and understanding the attacks first is a real neighbour and not a polite cross link.

Putting a model inside your own software is the third rung, and the iceberg is the honest picture of it: the call to the model is a few lines, and almost none of what decides whether the thing works. Getting it, cleaning it, and finding out what it says sits underneath it, because the retrieval everyone argues about is mostly a data problem wearing a model's clothes. On the fourth rung, models learned from data rather than written by hand are the oldest material here and still what wins on a table of rows.

Making images, video, audio and text with a model shares the bottom rung, filed apart because its failures are a different species: not a wrong fact but a right looking one, a voice that is nearly somebody's.

The topic about work is the one people arrive at last and should arrive at second. The clearest worked case is next door, where generated answers at the top of a results page changed what a click is worth and being the result someone clicks, without paying for it has had to redo its arithmetic because of it.

Which leaves the question this page cannot answer: which rung your own work is on. Most people say the bottom one, and most are on the second by the time anything they built runs on a schedule unwatched. Deciding whether to trust what came back is where that gets settled, one output at a time.

Alongside this

Above this