As AI begins to tackle problems at the frontier of mathematics, its progress raises a bigger question: how should human expertise evolve when machines can participate in the creation of knowledge?
For most of the AI era, the debate about work has centred on automation. Which tasks can AI perform? Which jobs will change? Which professions are safe? And, inevitably, which ones might disappear? However, a gathering of some of the world’s leading mathematicians at OpenAI this month raises a rather more difficult question. What happens when AI doesn’t simply automate routine work, but begins to participate at the frontier of human knowledge?
According The Washington Post, mathematicians gathered at OpenAI’s San Francisco headquarters to explore the rapidly advancing mathematical capabilities of AI systems and what those capabilities could mean for mathematical research. The discussion wasn’t simply about whether AI can solve difficult equations. Researchers are increasingly exploring whether AI can help prove theorems, identify connections, generate conjectures and contribute to problems much closer to the boundaries of mathematical knowledge.
There are important qualifications here. Today’s systems are not autonomous mathematicians, their outputs still require verification, and impressive results on particular problems should not be mistaken for mastery of mathematics as a discipline. However, the direction of travel is significant enough to warrant attention, because if AI can make meaningful contributions to work that previously demanded exceptional human expertise, we need to ask not only what these systems can do, but what their growing capabilities mean for the people working alongside them.
When automation reaches the frontier
Much of our thinking about automation is built around a familiar assumption. Technology handles the routine work. Humans handle the difficult work. Machines process transactions while accountants interpret them. Software searches documents while lawyers exercise judgement. AI drafts content while humans provide creativity and direction. Expertise sits towards the top of the pyramid.
Mathematics complicates that picture. Advanced mathematical research is highly specialised intellectual work. It requires years of training, abstract reasoning, creativity and the ability to work on problems for which an answer may not even be known. If AI can make meaningful contributions there, the boundary between automation and expertise becomes less clear. That doesn’t mean mathematicians are about to disappear. Nor does progress in mathematics automatically tell us what will happen in medicine, engineering, law or any other profession. However, it does challenge the assumption that increasing task complexity necessarily protects work from automation.
AI isn’t simply moving upwards through a list of increasingly complicated tasks. In some fields, it may be beginning to enter the processes through which new knowledge is created.
What can these systems actually do?
This is where some restraint is necessary. There is an enormous difference between solving a difficult mathematical problem and being a mathematician. Research involves much more than producing correct outputs. It involves understanding a field, recognising meaningful problems, developing intuition, challenging assumptions, communicating ideas and deciding which avenues of investigation deserve years of attention. AI systems can also produce convincing but incorrect reasoning. Mathematical claims ultimately need to be proved and verified, regardless of whether the initial insight came from a human or a machine.
The question, therefore, isn’t whether AI has suddenly reproduced human mathematical expertise in its entirety. It hasn’t. The more useful question is what happens when AI becomes capable enough to perform increasingly valuable parts of expert work. That is a subtler development, but potentially a more consequential one.
Finding answers was never the whole job
There is an important distinction between solving a problem and deciding that a problem is worth solving. Mathematics isn’t simply a collection of unanswered questions waiting for somebody to fill in the blanks. Researchers decide which areas deserve attention. They notice relationships between ideas. They develop intuitions about which approaches might lead somewhere interesting. They recognise when a result is surprising, important or potentially connected to something much larger. In other words, expertise isn’t simply the ability to produce an answer. It is also the ability to understand the significance of the answer.
If the cost of generating possible solutions, proofs, hypotheses or approaches falls dramatically, that part of expertise may become more visible. A researcher could potentially explore far more possibilities with AI than they could alone. The scarce resource may increasingly become deciding where to look, what to trust and what deserves further investigation.
From production to orchestration?
We can already see versions of this shift elsewhere. Software engineers work with systems capable of generating code. Scientists use AI to search literature and suggest candidate molecules. Designers can produce multiple concepts in minutes. Analysts can interrogate large datasets conversationally. In each case, greater machine capability can change what the human spends time doing. Less time may be spent producing every individual output. More may be spent defining objectives, providing context, evaluating results and deciding what happens next.
It is tempting to call this human oversight, but “orchestration” may be closer to what is actually required. The expert establishes the objective, understands the constraints, evaluates competing possibilities and connects outputs to a broader body of knowledge. AI expands what can be explored. Human expertise helps determine what should be explored and what the results mean.
At least, that is one possible future.
We shouldn’t assume the boundary will stay there
There is a danger, however, in drawing a comfortable line around whatever AI currently struggles to do. We have done this before. Creativity was once presented as an obviously human boundary. Complex language was another. Programming and reasoning have both been treated similarly. As systems improve, those boundaries become harder to defend.
There is therefore no reason to assume that generating valuable questions, identifying promising research directions or evaluating significance will remain exclusively human capabilities. AI systems capable of analysing enormous bodies of knowledge may eventually identify relationships that no individual researcher could reasonably discover alone. That possibility should neither be dismissed nor treated as inevitable. We simply don’t know yet how far these capabilities will develop.
A responsible response to that uncertainty isn’t to predict the disappearance of expertise or to insist that human judgement will always remain supreme. It is to prepare for both the opportunities and the risks created as the boundary moves.
The risk of losing the expertise we need
There is another side to increasing AI capability that deserves attention. If experts gradually delegate more of the intellectual process to AI, what happens to the skills required to evaluate its work? As we know, this is not a new problem. Automation has always created the possibility of skill degradation. When a system reliably performs a task, humans naturally practise that task less often. However, the stakes become particularly interesting when the automated capability is reasoning itself.
A mathematician who receives an AI-generated proof still needs sufficient expertise to understand and challenge it. A doctor using an AI recommendation needs the knowledge required to recognise an implausible conclusion. An engineer relying on generated designs must still understand the principles determining whether those designs are safe. The danger isn’t simply that AI gets something wrong. It is that humans gradually lose the ability to recognise when it has.
What should education do?
This makes the implications for education particularly important. One response to increasingly capable AI might be to reduce the emphasis on knowledge. If machines can retrieve information, solve problems and generate sophisticated outputs, why require humans to learn those things themselves? There is an obvious attraction to that argument, but it contains a serious weakness.
You cannot reliably judge an answer you don’t understand. You cannot recognise an important result without understanding the field around it. You cannot formulate good questions without understanding the problem, and meaningful oversight becomes difficult if the person supervising the system lacks the expertise required to challenge it.
Education, therefore, faces a difficult balancing act. People need to learn how to work effectively with AI, but AI literacy cannot become a substitute for disciplinary knowledge. We may need both. That means teaching people how to use intelligent systems while deliberately preserving the foundational knowledge, critical reasoning and practical experience necessary to evaluate what those systems produce.
Augmentation needs to be designed
The optimistic scenario is compelling. AI could allow researchers to explore more hypotheses, test ideas more quickly, discover connections across enormous bodies of literature and spend more time on the parts of research that require deeper judgement. Scientific and mathematical progress could accelerate, but augmentation doesn’t happen automatically simply because humans and AI are placed in the same workflow.
It has to be designed.
Institutions adopting these systems will need to think about where AI should act independently, where human verification is required, how outputs are validated and which skills must remain actively practised. They will also need ways of measuring whether AI is genuinely improving the quality of expert work rather than simply increasing its speed. Productivity and progress are not always the same thing.
A different definition of expertise
The discussion taking place in mathematics may, therefore, be an early glimpse of a much broader transition. For decades, professional expertise has been demonstrated partly through production. The programmer writes the code. The lawyer constructs the argument. The researcher develops the proof. The analyst builds the model.
Generative AI increasingly separates expertise from production. Producing something may become easier. Knowing whether it is correct, whether it matters, whether it should be trusted and what should happen next may remain considerably harder. Perhaps, that makes human expertise more valuable. Perhaps, increasingly capable AI will eventually challenge those abilities too.
Most likely, the relationship will continue changing as both the technology and the way we use it develop. The task for organisations, educators and professional bodies is, therefore, not to decide whether humans or AI should “win”. It is to ensure that as AI becomes capable of participating in expert work, we continue developing people capable of understanding, directing and challenging it.
The most interesting question raised by AI’s progress in mathematics isn’t whether machines can find the answer. It’s what happens to human expertise when they can.





