The AI Download #017

March 21st, 2025

Dear Reader,

Have you ever wondered how we (as a species) determine if a machine is truly “intelligent”? Long before ChatGPT entered our daily lives, notable mathematician Alan Turing was already thinking about this question. His simple yet profound test has shaped how we evaluate AI for over 70 years—and as these systems grow more capable, the ways we measure them continue to evolve in fascinating ways.

The Original Question: Can Machines Think?

In 1950, Alan Turing published a paper titled “Computing Machinery and Intelligence” that began with a deceptively simple question: “Can machines think?” Rather than getting lost in philosophical debates about the nature of consciousness, Turing proposed a practical test.

Imagine you’re texting with someone. You can’t see them or hear their voice–you only have their written responses. Could you tell if you were chatting with a human or a computer program? If you couldn’t reliably distinguish between the two, Turing suggested, then perhaps the machine deserves to be called “intelligent” in some meaningful way.

This became known as the Turing Test, and it’s remarkably similar to how many of us now interact with AI assistants daily. When you ask a question and receive a helpful, nuanced response, does it matter whether a human or AI composed it? Turing’s insight was that intelligence might be better judged by behaviour than by mechanism.

Beyond Pass/Fail: How the Measurement of AI Has Evolved

While elegant, the Turing Test has limitations. The binary pass/fail nature of the Turing Test doesn’t capture the spectrum of capabilities modern AI systems possess.

Today’s approach to evaluating AI has become much more nuanced:

  • Task-specific benchmarks measure performance on everything from grammar checking to medical diagnosis
  • Reasoning assessments evaluate whether AI can follow logical steps to solve problems
  • Creative tasks test if AI can generate novel, valuable outputs
  • Safety evaluations determine if AI systems can avoid harmful outputs when prompted

For example, when researchers want to measure an AI’s understanding of physics, they might present it with puzzles about objects in motion rather than asking it to fool a human judge in conversation. This gives us a more detailed picture of where these systems excel and where they still fall short.

Humanity’s Last Exam: A New Framework for the AI Era

What if we’re creating systems that will eventually surpass human capabilities across all domains?

In recent years, a provocative idea has emerged in AI research circles: what if we’re creating systems that will eventually surpass human capabilities across all domains? This concept, sometimes called “Humanity’s Last Exam,” suggests that the tests we design for AI today might be the last meaningful challenge humans pose before super-intelligent systems begin creating their own benchmarks.

Think about what this means for a moment. The math problems, coding challenges, and reasoning tests we’re using to evaluate today’s AI could be the final exams humans give to machines before they graduate beyond our level of intelligence.

These evaluations aren’t just academic exercises–they’re how we ensure AI systems align with human values before they potentially surpass human capabilities.

Measuring What Matters: Beyond Intelligence to Alignment

Perhaps the most important evolution in how we evaluate AI isn’t about intelligence at all–it’s about alignment with human values and goals.

When selecting an AI system to assist with tasks, raw performance isn’t the only concern. Teams need to know the system will protect privacy, provide unbiased recommendations, and explain its reasoning in understandable ways.

This reflects a broader shift in AI measurement. While we still care about capability, researchers are developing increasingly sophisticated ways to evaluate:

  • How well AI systems understand and respect human intent
  • Whether they can explain their reasoning in understandable terms
  • If they make fair and unbiased decisions
  • How they handle edge cases and uncertain situations

What’s the next step in evolution for intelligent machines?

The Tests We Create Reveal What We Value

There’s a fascinating aspect to this evolution in AI measurement that often goes unnoticed: the tests we design reveal what we truly value in intelligence.

When early AI researchers focused exclusively on logic puzzles and chess, they were expressing a particular view of intelligence centred on calculation and strategic thinking. As our evaluations expanded to include emotional intelligence, creativity, and ethical reasoning, we acknowledged a broader understanding of what makes intelligence valuable.

How we test AI systems today will influence what abilities those systems prioritise tomorrow. It’s like education–if we only test for memorisation, that’s what students will focus on developing.

What You Can Do: Becoming an Informed AI Evaluator

As AI becomes more integrated into our daily lives, each of us becomes an informal evaluator. Here are some practical questions you can ask when interacting with AI systems:

  1. Does it understand the nuance in my request, or am I having to oversimplify?
  2. When it makes a mistake, can it learn from the feedback I provide?
  3. Does it respect boundaries I set, or does it require me to repeatedly reinforce them?
  4. Can it explain its recommendations in terms that help me make better decisions?

These questions aren’t just academic–they help you determine which AI tools genuinely enhance your work and life, and which ones aren’t quite ready for prime time.

The Conversation Continues

The way we measure AI capabilities continues to evolve, reflecting our deepening understanding of both intelligence and what we want from our technological creations.

From Turing’s simple imitation game to today’s multifaceted evaluation frameworks, the question has expanded from “Can machines think?” to “Can machines think in ways that are beneficial, safe, and aligned with human flourishing?

I’d love to hear your thoughts on this topic. Have you found yourself evaluating AI systems in your work or personal life? What criteria matter most to you?


🛠️ Get My Agents for Creators, Builders and Doers.

My AI Creator’s Toolkit is a growing collection of lightweight, task-focused AI agents designed to help you get small jobs done faster. No fluff, no jargon, and no need to learn prompt engineering (but if you wanted to, there’s also a tool for that!).

Here are a few of the tools included:

  • Jeannie – helps you come up with ideas for what kind of GPTs you could build
  • Jamie – a YouTube assistant that helps you plan, script, and optimise your videos
  • Webster – turns your meeting notes into clear actions and takeaways
  • Seymour – quickly generates SEO-friendly meta descriptions from any webpage
  • Prompto – takes a messy idea and turns it into a clean, usable GPT prompt

The AI Creator’s Toolkit is all about cutting out repetitive tasks and giving you a shortcut to useful results. More tools are being added every month, so sign up today!

Check out the AI Toolkit

👉 You can try it for free (includes 10 credits)

That’s all for this week. See you next time!

Jim

linkedinexternal-linkexternal-linkmedium

Made with ❤️ in Valencia by Jim Christian. For feedback, please reach out to [email protected].

Built with Kit