Practical AI and future skills

Fei-Fei Li: Ask What an AI Benchmark Measures

When an AI system has an impressive score, how can I tell whether that score answers my question?

Self Growth Lessons
Choose, practice, reflect

A practical learning exercise

The Lesson

A score answers a question defined by a test. Before treating it as a reason to choose an AI tool, find out what that question was.

Fei-Fei Li is one of the authors of the ImageNet Large Scale Visual Recognition Challenge paper. The paper distinguishes tasks such as classifying an image and locating objects within it. Its classification discussion also distinguishes top-1 scoring, which checks the highest-ranked label, from top-5 scoring, which checks whether the reference label appears among up to five predictions. These are different evaluation rules for different questions; they are not interchangeable descriptions of a system’s result.

That distinction gives us a useful reading habit: ask what the system had to return and what counted as correct. A test that accepts several possible answers does not ask exactly the same thing as a task requiring one choice.

Imagine a fictional organizer sorting workshop questions into “date,” “materials” or “registration.” If an assistant lists all three categories for every message, the intended category will always appear somewhere in its list. That tells the organizer little about where to send the message. The organizer may need one appropriate destination, or a clear indication that a question covers more than one subject.

This imaginary example is much smaller than ImageNet and does not reproduce its test. It illustrates why you should read the rule behind a number. You need to know whether the test rewards one answer, a short list, an exact match or a different kind of result.

Also examine what the test leaves out. Sorting a message into the intended category does not establish that the reply is accurate, that a confusing message is handled well or that the organizer can correct a mistake. Those are separate questions. You can write them down rather than silently treating one score as evidence for all of them.

For your own small trials, define the task before looking at the result. Keep the examples and scoring rule visible. A useful record can be as simple as a list of sample inputs, the expected outputs, the returned outputs and the reasons for any disagreements.

Reflection

  • What action would I take because of this score?
  • Does the test require the kind of output my task needs?
  • Which examples were included, and which situations are missing?
  • What important question would still be unanswered by a perfect result?

Practice

Original SelfGrowthVideos exercise: score the same fictional answers two ways. This worksheet is not ImageNet, an actual model evaluation or Fei-Fei Li’s personal method.

Use the three categories date, materials and registration. For this exercise, each message has one assigned reference category.

Fictional messageReference categoryReturned list, in order
What day is the workshop?datedate, registration, materials
Do I need my own pencils?materialsregistration, materials, date
Where do I sign up?registrationmaterials, date, registration
What time should I arrive?datematerials, date, registration
  1. Score the first answer only. Mark a row correct only if the first returned category matches the reference. One of the four rows passes.
  2. Score inclusion anywhere in this three-item list. Mark a row correct if the reference category appears anywhere. All four rows pass.
  3. Explain the difference in words. The returned lists did not change. The scoring rule did. Neither result proves that these answers can route the organizer’s messages usefully.
  4. Choose a practical requirement. Suppose the organizer needs one destination per message. Write that requirement separately from the inclusion score.
  5. Add one harder input. Invent a message asking both when to arrive and what to bring. Decide how you would handle multiple subjects before judging an answer to it.
  6. Record one missing test. For example: “Can the organizer see and correct a wrong destination?” Describe an observation that would help answer it.

Keep the four-row example and your chosen rule together. If you change the rule, label the change so someone else can understand why a later result differs.

Review

When you next read an AI comparison, try to identify the task, sample inputs, allowed outputs and scoring rule. If the source does not provide them, record that uncertainty instead of guessing.

Ask whether the result supports your intended use. Choose one additional question to investigate, rather than trying to make a single number stand for the whole experience.

Go Deeper

Sister brand · Vacation Club Promo

All-inclusive resort stays from $435

Qualified couples: promotional Mexico and Caribbean packages at real resorts. Short presentation during your stay — enjoy the rest of the trip.

Qualifications apply · presentation during stay · no purchase required

Subscribe YouTube Suggest