When did you last change your mind about what AI can do?
Not when did you last use ChatGPT. Not when did somebody show you an interesting demo.
When did you actually update your understanding of what artificial intelligence is capable of?
For many people, that view was formed surprisingly early.
Perhaps you experimented with ChatGPT after it appeared in late 2022. It wrote some impressive text, confidently got some things wrong and produced enough generic rubbish to make its limitations obvious.
Maybe you tried again when GPT-4 arrived.
Perhaps you’ve used Microsoft Copilot, Google’s Gemini or one of the growing number of AI products now appearing inside everyday software.
Each encounter creates a mental model.
This is what AI is good at.
This is where it struggles.
This could be useful.
This isn’t good enough.
That’s completely rational.
The problem is that the technology being judged isn’t standing still.
And if capability is improving quickly enough, a perfectly sensible conclusion can become a bad assumption without anybody noticing.
Eighteen months is a long time in AI#
Consider how much changed between the end of 2022 and the middle of 2024.
The first version of ChatGPT introduced millions of people to the idea that they could have a surprisingly sophisticated conversation with a machine.
Then came GPT-4 in March 2023.
OpenAI reported that GPT-4 was more reliable, creative and capable of handling more nuanced instructions than GPT-3.5. It could also work with images as inputs, rather than text alone.
One benchmark illustrated the size of the improvement. GPT-4 performed around the top 10% of test takers on a simulated bar examination. GPT-3.5 had performed around the bottom 10%.
That didn’t make GPT-4 a lawyer.
But it demonstrated something important: the capability wasn’t static.
Then the field broadened.
In December 2023, Google introduced Gemini, designed from the outset to work across text, images, audio, video and code.
In March 2024, Anthropic introduced its Claude 3 family, reporting substantial improvements across reasoning, mathematics, coding and visual understanding.
And then, on 13 May 2024, OpenAI demonstrated GPT-4o.
This time the change was particularly easy to understand without looking at a benchmark.
You could watch it.
The interface is changing again#
GPT-4o can work across text, vision and audio.
In OpenAI’s demonstrations, people spoke to the system, interrupted it, showed it things through a camera and received spoken responses with remarkably little delay.
OpenAI reported average audio response times of 320 milliseconds — approaching the response time of a human conversation.
The significance isn’t that we now have a better voice assistant.
It’s what happens when several capabilities begin to converge.
A machine that can process language is interesting.
A machine that can interpret images is interesting.
A machine that can understand speech is interesting.
A machine that can generate software is interesting.
Start combining those capabilities in a system that can move between them naturally and the range of potential applications expands.
The interaction changes too.
Typing a carefully constructed prompt into a text box still feels like using a piece of software.
Talking naturally to something that can hear what you’re saying, see what you’re looking at and respond almost immediately begins to feel different.
There is plenty that today’s systems still cannot do reliably.
But compare the experience being demonstrated now with the ChatGPT that appeared eighteen months ago.
The direction is difficult to miss.
The danger isn’t simply underestimating what AI can do today. It’s assuming today’s limitations will remain limitations for long enough to build a strategy around them.
Business assumptions have expiry dates#
Businesses make technology decisions all the time.
Usually, an assessment has a reasonably useful shelf life.
If a piece of software can’t perform an important function today, there is a good chance it still won’t perform it next month.
AI creates a different problem if the underlying capability is moving rapidly.
Imagine a company investigates whether AI could handle a particular customer-service workflow in early 2023.
The conclusion is no.
The answers aren’t reliable enough. The system struggles with the company’s information. The experience isn’t good enough for customers.
That might be exactly the right decision.
But what happens next?
Frequently, the conclusion survives longer than the evidence supporting it.
The organisation doesn’t remember:
The technology wasn’t capable enough when we evaluated it in February 2023.
It remembers:
AI can’t do that.
Those are very different statements.
The first is an assessment made at a particular point in time.
The second becomes organisational knowledge.
And organisational knowledge is sticky.
People repeat it in meetings. Investment decisions reflect it. Projects don’t get proposed because everybody already “knows” they won’t work.
Meanwhile, the technology that produced the original conclusion changes.
In a rapidly improving technology, “not possible” can have a surprisingly short shelf life.
Not every demo becomes a business#
There is an obvious counterargument to all this.
Technology companies have every incentive to make their latest systems look extraordinary.
A controlled demonstration isn’t a production environment.
An academic benchmark isn’t a customer.
And being able to perform a task once is very different from performing it accurately, securely and economically thousands of times inside a business.
These distinctions matter.
Today’s AI systems still hallucinate. They make unpredictable mistakes. They can misunderstand instructions. The consequences of an error vary enormously depending on whether you’re generating ideas for a marketing campaign or making a decision involving someone’s money, health or legal rights.
Integration matters.
Data matters.
Security matters.
Governance matters.
Economics matter.
People matter.
A spectacular model demonstration can therefore coexist with a terrible business case.
That isn’t an argument against taking the capability curve seriously.
It’s an argument for understanding what you’re actually evaluating.
The mistake would be jumping from AI is improving rapidly to therefore every AI application makes sense.
It doesn’t.
But the opposite mistake is equally dangerous: discovering that something doesn’t work today and treating that conclusion as permanent.
Better and cheaper can happen together#
There is another dimension to the capability curve that businesses shouldn’t ignore.
Quality isn’t the only thing changing.
The economics can change too.
When OpenAI introduced GPT-4o, it said the model matched GPT-4 Turbo’s performance on English text and code while operating faster and at half the API price.
Anthropic’s Claude 3 family similarly offered different trade-offs between intelligence, speed and cost, with its Haiku model explicitly designed as a faster, lower-cost option for high-volume workloads.
This matters because a business case can become viable in more than one way.
A model can become capable of doing something it previously couldn’t.
But an existing capability can also become fast enough or cheap enough to use at scale.
That’s a different kind of progress.
A business might correctly conclude today that using AI for a particular process would save £1 but cost £2.
If the cost of the underlying intelligence falls substantially, nobody needs to invent a new capability for that decision to change.
The technology simply has to cross the economic threshold.
AI opportunities can emerge because the models get smarter, but they can also emerge because existing intelligence becomes faster, cheaper and easier to deploy.
For a business, both curves matter.
Stop asking whether AI can do it#
This changes how I think businesses should evaluate AI opportunities.
The obvious question is:
Can AI do this?
But that increasingly feels incomplete.
A better set of questions might be:
Can AI do this reliably enough today?
At what cost?
What prevents it from being useful?
Is that constraint fundamental, or is it simply a limitation of the current technology?
When should we test it again?
That last question may become particularly important.
A rejected AI opportunity shouldn’t necessarily disappear into a PowerPoint deck.
Perhaps it needs an expiry date.
If the potential economic value is significant but the technology isn’t ready, record why it failed and revisit it when something material changes.
A better model arrives.
Costs fall.
Context windows expand.
Reliability improves.
A missing integration becomes available.
The organisation’s own data improves.
Then test the assumption again.
This doesn’t require businesses to chase every model release or reorganise themselves every time an AI company publishes a benchmark.
It requires something much simpler:
separating today’s limitations from permanent ones.
The most dangerous word may be “can’t”#
We don’t know how long the current rate of AI progress will continue.
It may slow.
Some problems will prove much harder than impressive demonstrations suggest. Scaling today’s techniques may eventually produce diminishing returns. Reliability may constrain adoption in areas where errors carry serious consequences.
There is no sensible reason to assume an exponential capability curve continues indefinitely.
But businesses don’t need to make that prediction.
They only need to recognise what has already happened.
In roughly eighteen months, the mainstream experience of generative AI has moved from exchanging text with an impressive but unreliable chatbot to interacting with systems that can increasingly reason across text, images, audio, video and code.
That is enough to change how technology assessments should be treated.
The next time someone says AI can’t perform a valuable task inside a business, the right response isn’t necessarily to disagree.
They may be completely right.
Just add one question:
When did we last check?
