About Platform Resources Articles Contact
AI & Business  ·  IMA AI

The model reviewer has
never deployed a model.

A model ships in the morning. By dinner, YouTube has expert verdicts on it — delivered by people who have never deployed a model inside a real business, not once. So what exactly qualifies them to rate anything, and why do we keep listening?

Published  August 2026
By  Chin Qi Yong, CEO — IMA AI
© 2026 Chin Qi Yong
Read time  ~5 min

Every model release now follows the same script. The provider drops a new model in the morning. Within hours, the review videos land: “I tested the new model — here's my HONEST verdict.”

Business owners forward these videos to me every week and ask the same question: should we switch? My honest answer is that the man in the video cannot know. He has never run the thing he is rating.

The overnight expert

Watch what the “test” actually is. A riddle. Counting letters in a word. One coding puzzle. “Write me a poem.” Then a scoring table the reviewer invented himself, weighted however he feels, and a confident conclusion: this changes everything — or this one flopped.

That is not a methodology. That is a man setting his own exam questions, marking his own answers, and calling the result a rating. No two reviewers test the same things. None publish an error rate. None re-test a month later, after the model has been patched twice. The measurement is whatever the reviewer thinks is right — and what he thinks is right is whatever fits in a twelve-minute video.

In what other field does this work?

Think about how advice works everywhere else. An auditor signs accounts because a licence and a liability sit behind the signature. A doctor recommends treatment because he has treated patients and answers for the outcome. Even a food reviewer eats the actual food.

The model reviewer is rating a deployment tool without ever deploying it. He has never wired a model into a live workflow, never watched it fail on a client deadline, never paid the invoice for its mistakes, never had staff waiting on its output at month-end. His entire contact with the product is the demo — the test drive in the parking lot. Then he rates the lorry without ever loading cargo.

I have written about this species before — the conference speaker whose entire AI expertise was the free tier. Same market, new stage. The stage moved to YouTube, and the audience got bigger.

Why they rate anyway

So why do people with no deployments and no methodology hand out verdicts? Because the verdict is not analysis. The verdict is the product.

The incentives are simple. Speed: the video must be live while the model trends — which is exactly the window in which nothing real can be known yet. Confidence: “I'm not sure, ask me in a month” gets no clicks; “this kills the competition” does. And zero consequence: when the verdict turns out wrong, nothing happens to the reviewer. The viewer who switched his business tools on that advice pays the bill. The reviewer uploads the next review.

Advice without consequences is not advice — it is content. I made the same point about the course sellers: the test is always operator versus guru. Does this person run what he teaches? Ask a model reviewer one question — what did this model break in your business last month? An operator has a story. A reviewer has a thumbnail.

And before anyone answers “trust the official benchmarks instead” — the industry's own shared scoreboard was gamed by Meta, whose special leaderboard build of Llama 4 ranked #2 while the version you could download ranked 32nd. Researchers separately found roughly 6.5% of MMLU's exam questions are simply marked wrong. That is the professional version of testing. The homemade YouTube version does not even have an answer key.

What testing looks like when you carry the consequences

Here is the difference deployment makes. We generate images for client work every day. The launch reviews scored the image models the usual way — cinematic demo prompts, side-by-side beauty contests. My staff put one of the leading models into the real pipeline and found the defect within days: it cannot render text reliably. Fonts break. Words come out wrong.

Not one review mentioned it — because it only shows up when someone has to ship the output to a paying client. Our whole production pipeline now builds text as a separate overlay layer, because of a flaw no reviewer ever found.

That is what a real rating is: what it cost, what it got wrong, and how many hours a human spent repairing the output. Three numbers, all of them living inside your business flow, none of them visible from a demo. It is the same way we chose our own models — benchmarks and price built the shortlist, but the decision came from running our actual work: roughly 57x cheaper for a single-digit gap, proven on our tasks, not on someone's riddle.

To be fair about what creators are for

Review videos have one legitimate job: discovery. They tell you a model exists and what it claims. That is news, and news is fine. Some creators are honest about being entertainers, and entertainment is fine too.

The failure sits on both sides of the screen. The reviewer claims an authority he never earned. And the viewer accepts procurement advice from someone who carries none of the consequences. If you make business decisions from launch-week videos, the reviewer is not the only one skipping the real test.

The bottom line

Run your own review. Pick one real task from your actual workflow — not a riddle, the work itself. Run the candidate model on it for a week inside the live flow. Measure three things: the cost, the error rate, the rework hours. Switch on those numbers or do not switch at all.

The reviewer's verdict was out before the model finished downloading. Yours can afford to take a week — it is the only verdict that gets graded, and the examiner is your P&L, not the algorithm.

The bottom line
Advice is worth exactly the consequences the adviser carries. The reviewer risks nothing, so his verdict is worth nothing — the only benchmark that cannot be gamed is your own invoice.
CQ
Chin Qi Yong
CEO, IMA AI
Chin Qi Yong is the CEO of IMA AI — building the infrastructure layer for agent-era commerce and identity in Malaysia. IMA AI's products are designed for the world where AI agents transact, verify, and operate on behalf of humans.
Follow on LinkedIn

Published by IMA AI — August 2026. Written by a CEO who tests models on his own payroll, not on his own channel.