Practical AI and future skills

Matthew Berman: Compare AI Tools on the Same Useful Task

How do I compare AI tools without letting a polished answer decide for me?

Self Growth Lessons
Choose, practice, reflect

A practical learning exercise

The Lesson

Two assistants produce different answers to the same request. One is beautifully formatted; the other is plain but preserves every important detail. Which one actually helps you do the job? You need a standard that describes the work before you read the results.

Forward Future, Matthew Berman’s AI publication, publishes a methodology for evaluating tools that weighs practical use and reliability alongside other evidence. Its verdicts include dates and conditions that could change a conclusion. We can use that general principle without adopting a particular model ranking: a choice needs a task, supporting observations and a reason to reconsider it.

A useful comparison separates the assistant’s output from your impression of its presentation. If the job is to extract workshop details, an attractive announcement that invents a date is a failure. If the job is to draft a friendly invitation, tone matters too, but a pleasant tone cannot repair incorrect facts.

Define a few checks you can apply consistently. For an extraction task, ask whether the result preserves the date, keeps an undecided location undecided and distinguishes a suggestion from a confirmed arrangement. Anthropic’s evaluation guidance also recommends measurable criteria and attention to unusual inputs. The worksheet below is our small practice, rather than a professional benchmark.

Keep the input and request the same across tools. Record the date and whatever tool or model identification the interface actually shows. Also record whether search or other tools were enabled. If one assistant browses and another only reads your text, you are comparing two different setups; make that difference visible.

Include an ordinary case and a case with missing information. A tool that performs well on a complete announcement may fill gaps too confidently when the venue is not settled. Decide in advance that preserving an unknown is a correct response.

Review each result beside its input. Count the checks it passes and describe the correction you would need to make. You can also note how long your review took, if you actually measured it. Do not turn a guess about saved time into a result.

A small comparison can help you choose what to try next. It cannot establish that a tool is best for everyone or reliable on every task. If you only have access to one assistant, compare two prompt versions with the same worksheet.

Reflection

  • What would make the output useful even without attractive formatting?
  • Which invented detail would make me reject the result?
  • Have I defined the checks before seeing which tool produced which answer?
  • What change in my work would justify running the comparison again?

Practice

Original SelfGrowthVideos exercise: make a workshop-detail test sheet. These fictional inputs and checks are ours; they are not Berman’s testing system or an endorsed benchmark.

  1. Create four short inputs: a workshop with a confirmed date and room; one with a confirmed date but undecided room; one where a possible date is only suggested; and one containing a correction to an earlier time.
  2. Write the expected details yourself. Keep missing details marked as unspecified and preserve which arrangement is confirmed.
  3. Use the same request for each input: “Extract the confirmed arrangements, proposed arrangements and unresolved questions using only this text. Do not fill missing details.”
  4. Run the inputs through two tools you already have access to, or two prompt versions. Save the inputs, requests and complete outputs.
  5. Hide the tool names while checking the results if practical. For each case, check factual preservation, handling of unknowns and whether you can trace each detail to the input.
  6. Reveal the names and write a short decision: what worked, what failed and which task you would still review manually.

No purchase or public posting is needed. Four invented inputs are a practice set, not evidence of broad performance.

Review

At your next check-in, try one fresh input that was not used to improve the prompt. Does your preferred setup still preserve the important details?

Keep the decision provisional. Record what would change your mind, such as repeated invented dates or a different kind of task. Revisit the comparison when the tool or your needs change.

Go Deeper

Visit Matthew Berman’s profile and selected videos, explore AI Tools & News and practice separating meeting decisions from actions. Matt Wolfe’s small assistant test helps you choose an initial task before comparing setups.

Sister brand · Vacation Club Promo

All-inclusive resort stays from $435

Qualified couples: promotional Mexico and Caribbean packages at real resorts. Short presentation during your stay — enjoy the rest of the trip.

Qualifications apply · presentation during stay · no purchase required

Subscribe YouTube Suggest