中文
Search the unfinished book

Enter a keyword to search published articles.

← Back to articles
People and AI

GPT Added One Number to My Grill-Me Exercise

A confidence score from 0 to 100 revealed something a correct answer alone could not: the learner's sense of how well they understood it.

A learning report with no obvious flaws

Recently, I gave a newcomer with almost no prior experience an introduction to Codex. Afterward, I asked him to write a learning report about what he had taken away.

To be honest, the report was impressively well-formed. It covered all the expected concepts, had a complete structure, and offered little to fault. That was precisely why I struggled to tell how much he had really understood.

I did not want to quiz him point by point from the report or simply ask, “How well did you learn it?” Instead, I started talking with GPT about a practice called grill-me: follow a person’s answers with question after question, peeling back vague judgments and unspoken assumptions until they have to explain what they mean.

Could GPT use that idea to build a small exercise quickly? I was not looking for a grade, only a way to help the newcomer see his actual level and room to improve.

The number GPT unexpectedly added

The project was running in less than ten minutes.

It kept the idea of persistent questioning. It introduced an unfamiliar situation, waited for the learner’s response, then pressed further along his line of thought: What was the basis for his answer? What would change under different conditions? What would he do next? It offered no model answer and made it difficult to recite the training material by rote.

But beyond those expected settings, GPT added a requirement to every response:

Also state how confident you are in this answer, from 0 to 100.

At first I barely noticed it. I sent the exercise to the newcomer as it was.

After a few rounds, he asked, “Doesn’t this subjective number mean very little? It isn’t part of a grade. I can put down whatever feels right.”

Only then did I really stop to think about that number, labeled confidence.

His objection was fair. Writing 80 does not mean a person has achieved a score of 80. Writing 100 certainly does not make an incorrect answer true.

A subjective number is not a measure of achievement.

A second axis beside the answer

Traditional tests often leave us with a binary result: right or wrong.

Reality is more complicated. The same correct answer might come from deep understanding, a memorized conclusion that remains hazy, or a lucky guess. Yet the final score records all three as the same unblemished right.

That word makes it easy to stop: I got it right; I know it; I have mastered it.

A wrong answer may expose ignorance, but a right one can hide it more effectively.

Attach a confidence number to the answer, however, and right is no longer enough. If you were correct but only 60 percent confident, where did the other 40 percent go?

The number becomes uncomfortable. It prevents you from letting a fortunate correct answer cover a gap in your understanding. This was the second axis GPT had added to the exercise:

The answer responds to the question. Confidence asks how much you yourself believe your answer.

Once the two axes meet, one correct answer does not automatically mean mastery.

One hundred is a coordinate, not a perfect score

Cross correctness with confidence and the assessment is no longer a simple black-and-white plane.

One combination deserves particular attention: a completely wrong answer paired with absolute confidence. This is more than a gap in knowledge. It is a gap the learner cannot even see. On their map the road looks complete; in reality it has already fallen away. A confidently wrong answer marked 100 exposes the illusion of complete understanding.

Put right and wrong on one axis and high and low confidence on the other, and four situations emerge:

  • Right answer, high confidence: understanding and self-assessment are most closely aligned.
  • Right answer, low confidence: being correct has not removed uncertainty. Was it knowledge or luck? That remains to be tested.
  • Wrong answer, low confidence: this is not necessarily alarming. The learner can already see that something is unknown.
  • Wrong answer, high confidence: the most dangerous case—not knowing that you do not know, as in the wrong answer marked 100.

In this picture, 100 is not a perfect score and 30 is not a failing grade. Without an answer to anchor them, both are just subjective feelings. Alongside an answer, they help locate the gap between a person’s judgment and the result.

More importantly, the number must be allowed to move.

Under persistent questioning, an initial confidence of 90 might fall to 50. That is not a setback. It can be the process of finding your actual position. Conversely, if evidence has overturned an answer but confidence remains fixed at 100, what has stopped moving is the ability to update one’s understanding.

Learning is not simply pushing confidence from 60 to 100. It is knowing how confidence should change when new evidence arrives.

I finally understood why GPT’s extra number mattered. It was never a score for ranking learners. It cannot tell you how much of the world you know. It can show how far your confidence stands from what you actually understand.

Translation: Codex prepared this English version from Biaoo’s published Chinese article. Biaoo wrote the original account and its interpretation.

WECHAT · 微信公众号QR code for 仓颉的未完书 on WeChat仓颉的未完书

Scan with WeChat to follow.