Use of Generative AI in solving mathematical problems
Computing Project
Author: Sam Seifi
Abstract
Generative Artificial Intelligence (GAI) is a specific type of Artificial Intelligence (AI) where the program is trained from large sets of data to performing tasks; these tasks can include creating and solving math problems. While GAI can be used appropriately to help augment our creativity and productivity. It can also limit creativity and cognition, which will depreciate the value of being able to problem solve where for solving math problems is a core fundamental that is used throughout the subject.
How does KCL’s policies reflect these strengths and weaknesses?
King’s College London (KCL) (and the Russell Group universities as a whole) have produced policies that reflect the increased use of AI in wider society by encouraging its use in their institutions. In particular, the ‘Russell Group’s five principles’ convey that Group’s desire to help students and teachers alike to become more ‘AI literate’ while still ‘ensuring academic rigour and integrity is upheld’.
At the student level, KCL advises the use of generative AI mainly as an advanced word processing software – to summarise, transcribe, reformat or translate text or speech. Additionally, as is suggested in Chloe’s section, it advocates for students to use the technology (ChatGPT was used as an example) as a dialogic tutor, where the student can ask follow-up questions to a query and get immediate, tailored responses. However, there is no [easily accessible] advice relating to its use in mathematical problems and therefore a major gap in guidance on the potential pitfalls outlined in Sam’s section, which may lead students into incorrectly believing an AI’s solution as gospel/over-relying on its ability to try (and fail) questions requiring lateral instead of logical thinking.
At the departmental level, KCL generally advises against ‘integrating generative AI into most summative assessments. The guidance is for educators to engage with the technology themselves and base examinations around skills generative AIs cannot easily apply (e.g. extended logical thinking), though it does state a point which is reiterated in Stefan’s section, being that the advances in this technology will rapidly outpace short-term safeguards put in place by individual teachers. However, it does not suggest any long-term departmental changes aside from more communication between staff members (a suggestion that cannot be credited as novel).
Overall, though KCL’s guidance promotes generative AI’s use in secretarial tasks (i.e. word processing), its guidance lacks clarity on mathematical applications, leaving students vulnerable to over-reliance.
How can over-reliance on AI as a study aid be detrimental?
AI has become increasingly more widely used, especially with the recent generation of students as study assistants, particularly when trying to tackle something new. As a matter of fact, Pew Research Centre suggests that 79% of US adults interact with AI in some way at least once a day. With the rise of the use of AI it is important to assess the detriment that it can cause to the users and to their education.
(Zhai, 2024) delves into this, with the main argument stating that regular utilization of these systems results in an overall decline in cognitive abilities, including but not limited to; an overall decrease in the capacity for information retention, ignoring/ and minimizing the effects of critical ethical concerns such as: privacy, plagiarism, acceptance and sharing of misleading content. Critical thinking, decision making, and analytical thinking are incredibly vital among students, especially those in higher education, but also those younger who are beginning to cultivate their problem-solving skills.
Studies conducted on the use of Generative AI show a newfound over reliance on tools such as ChatGPT, Gao et al (2023) proving a troubling trend where users showed an over acceptance of AI output and hallucinations without validation or inquiry. As stated in the study 32% of blind human reviewers believed that AI generated work was real, this overdependence is furthermore intensified by cognitive biases where judgment is deviated from rationality and interest to the use of mental shortcuts, leading to uncritical acceptance of AI-generated information all due to laziness.
What are the general benefits of using generative AI in the learning process?
Successfully incorporating generative AI into education presents a distinct set of both significant benefits and challenges. For example, it should be emphasised that it can deepen student understanding by acting as a form of intellectual scaffolding – by generating both verbal and visual explanations, examples and multiple perspectives on a certain topic, AI helps learners clarify complex ideas and engage more effectively with material. This kind of support may encourage students not to rely on AI for answers, but to use it as a tool to refine their reasoning and strengthen their conceptual grasp.
In addition to support understanding, generative AI is highly valuable in personalisation. It can adapt explanations, difficulty levels, and feedback to meet individual learner needs, allowing students to progress at their own pace and revisit challenging areas. This makes learning more accessible and responsive, particularly for students who may require additional support. Furthermore, AI can provide immediate feedback and generate tailored learning materials, helping teachers deliver more effective instruction without adding to their workload. A study was conducted, (Jussi S. Jauhiainen, 2024), between pupils aged 8-14, it was found that AI-generated content aligned well with curriculum goals and improved student motivation, as many learners enjoyed the tailored materials and felt they better supported their understanding.
Overall, generative AI can enhance education by helping students understand difficult ideas through clear explanations, examples, while also personalising learning to each student’s needs. It offers adaptive support, immediate feedback, and tailored materials that improve engagement and accessibility, with research suggesting that AI generated content can boost motivation and align well with curriculum goals.
“AI has been unreliable in solving mathematics problems”
Throughout the evolution of Artificial Intelligence (AI), making a model to outperform the next has always been the norm. The performance of math questions for AI has also been excelled at a exponential rate to where even modern AI such at ChatGPT and Gemini can solve difficult math problems at the level of graduate to research level maths. This has in turn introduced a reliance on AI where people seemingly think it can solve undergraduate or lower education level solutions at a level where there isn’t human validation, though this is not the case given that even the most successful model (Gemini 3 Pro) only got ~37% of the questions correct when looking at AI model performance in figure 1 (FrontierMath, 2025). The figure shows the variation between different kinds of LLMs and their performance with questions that they haven’t been trained on before.
Furthermore, a study by (Iman Mirzadeh, 2024) suggests that the LLM aren’t really “solving” the problem at hand but rather predicting the optimal outcome to satisfy the user’s needs and that outputs are optimized for coherence rather than correctness, this makes errors often appear in the same confidence level as accurate solution making them difficult to verify. As a result, they can produce answers that sound rigorous while containing logical and arithmetic mistakes in cases of misuses of theorems and more.
For example, when testing LLMs with given data and changing numerical values or adding additional steps to solving the same problem, the performance of the LLMs result and generation of solution declined increasingly. This indicates that LLMs are sensitive to changes that may involve any sort of reasoning capacity limiting its capabilities to solving math problems.
Similarly, Google’s Minerva model, although having a strong record (fig 3) on quantitative reasoning tasks, reportedly had an equal number of calculation and reasoning errors which notably came though only “false positives”, which is arriving at the solution through flawed logic (Lewkowycz, 2022). This poses an even bigger issue, if a LLM can solve a problem yet it cannot explain the solution without produce flawed logic then there is no purpose for a user to learn from it, which will make it unreliable for users to learn from.
In conclusion, evidence shows that LLMs don’t reliably control logical deductions which is a non-negotiable layer of mathematical reasoning, yet LLMs exploit these patterns of problem-solving using training data and probability heuristics to make the outputs of their solutions sound coherent and ‘true’. Yet subtle errors and logical flaws that and typically ignored by the student can be taken for granted in the level of reliability that is given to the AI.
What types of mathematical problems does generative AI answer well?
To provide a conclusive problem type that AI can solve well, it must first be understood that generative AI’s performance in mathematical problems is dependent on multiple factors including the specific model used, therefore, its intended function must be taken into consideration.
To show this, conducted an experiment where I gave the same mathematical problems to 2 different generative AI models: Google Gemini, a ‘general-purpose’ model that is usually used for summarising google search results. Gauth AI, A model that was created with the specific intent of solving math problems and providing step-by-step solutions. These models were chosen as they have differing intended functions and so comparing them should give a comprehensive understanding on generative AI’s ability to solve mathematical problems.
I used 3 problems that test various mathematical concepts. Due to AI using an element of randomness in its ‘decisions and reasoning’, different solutions can be produced even if the exact same problem is given. Due to this, each problem is given 5 times. Both AI models used take input in the form of an image uploaded so this is the method used. I ensured that the input was clear so that incorrect input or errors with the image was not a cause of potential error. These are the results of my experiment:
From my experimental research (Figure 4) it can be deduced that generative AI is most effective at solving problems with symbolic algebraic manipulation and struggles with problems that require analysing diagrams (finding missing angles or applying circle theorems) or using graphs such as finding a line of best fit. However, this is only truly reliable if it verifies it’s results with a non-AI source. The conclusion of this is that while generative AI’s ability to solve problems have greatly improved in recent years, it is still incapable of reliably giving a correct answer without verifying with non-AI sources such as WolframAlpha.
AI’s mathematical ability will only increase as more specialised models are created, this is shown in the below image where the OpenMath – Nemotron 32B model achieved a 76.6% accuracy in mathematical problems (MarkTechPost – NVIDIA - AI):
The main problems that generative AI currently excels at is the handling of big data. It is fundamentally difficult to handle big data using standard systems due to its volume (amount of data), velocity (rate at which data is generated or collected) and variety (many different data types to be stored). Examples of this would be storing posts uploaded to a social media platform. Generative AI has the advantage of being able to ‘adapt’ so it can find patterns and form outputs from big data. Using generative AI, companies can now create giant big-data stores that can be easily accessed as the chatbot style interface allows for a more user-friendly experience than traditional database interfaces such as SQL (SeaGate, 2025).
Generative AI is also capable of making decisions based on large amounts of data at a much higher speed than a human can. An application of this is AI in finance, taking in huge amounts of data and recognising patterns that signal security risks such as fraud or even to automate stock trading. (Intel, 2025). Most currently available models (Gemini, ChatGPT, etc) are incapable of solving math problems effectively because they are trained as large language models, so their main function is to predict the next word in text, not solve mathematical problems. The reason it can handle large amounts of data but not math problems is because generative AI relies on pattern recognition which cannot always be applied effectively to these problems.
To summarise, while most generative AI models currently available do not have a capability to reliably solve mathematical problems that a student may encounter, its ability to form correct solutions to these problems is rapidly increasing as shown in (Nvidea, 2025). The current main application of generative AI in mathematics is to handle large amounts of data on scales that would be impossible for humans to handle and process. AI is also used as an interface so that humans can have easy access to databases.
How would we want to use AI in our studies:
Currently AI can be seen as a very powerful study tool, a personal tutor where you can ask a 1 to 1 question and get a response almost instantaneously rather than for example emailing your lecturer about a certain topic and waiting possibly hours for a response.
Some of the many ways we believe AI should be utilised is by asking it to summarise lecture notes, explaining ideas behind complex mathematical equations, checking your working out and generating specific questions for certain topics.
As much as this is useful, it cannot be taken as a main source of revision as of course AI can make mistakes. Hence, it is always important to fact check AI generated answers with textbooks or lecture notes or any other reputable sources. In addition to this, questions should always be attempted before asking AI for feedback. Furthermore, AI should be always considered as a second opinion, not primary.
While the use of AI as a study tool for solving mathematical problems and understanding mathematical concepts is great, it must never be taken to such an extent to which a student starts to lack academic integrity. Whether it be heavy reliance of AI by asking it to answer mathematical problems without trying themselves, copying AI generated solutions into assessments or asking AI to solve graded questions.
To conclude, KCL’s AI policy clearly defines the ways in which students could utilise generative AI in their courses while still having academic integrity. KCL encourages the use of generative AI to help students advance their understanding of concepts and solving capabilities.
References
Giannakos, M. (2024) The promise and challenges of generative AI in education. Behaviour & Information Technology. https://doi.org/10.1080/0144929X.2024.2394886
Elliot Glazer, E. E.-S. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. https://arxiv.org/abs/2411.04872.
FrontierMath. (2025). Benchmarking AI against advanced mathematical research. AI Model Performance on FrontierMath.
Gao, C. H. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. https://doi.org/10.1038/s41746-023-00819-6.
Iman Mirzadeh, K. A. (2024). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. https://arxiv.org/abs/2410.05229.
Intel. (2025). Artificial Intelligence (AI) Use Cases and Applications. intel.
Jussi S. Jauhiainen, A. G. (2024). Generative AI and education: dynamic personalization of pupils’ school learning material with ChatGPT. Sec. Digital Learning Innovations.
Lewkowycz, A. A.-S.-A. (2022). Solving Quantitative Reasoning Problems with Language Models. NeurIPS. https://research.google/blog/minerva-solving-quantitative-reasoning-problems-with-language-models.
Nvidea. (2025). NVIDIA AI Releases OpenMath-Nemotron-32B and 14B-Kaggle: Advanced AI Models for Mathematical Reasoning that Secured Firs. marktechpost.
SeaGate. (2025). Generative AI is finally enabling the promise of big data. https://www.seagate.com/gb/en/stories/articles/generative-ai-is-finally-enabling-the-promise-of-big-data.
Zhai, C. W. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: a systematic review. https://doi.org/10.1186/s40561-024-00316-7.
King’s College London (2023).
- Macro-level guidance: University-wide principles and policy https://www.kcl.ac.uk/about/strategy/learning-and-teaching/ai-guidance/macro-level
- Generative AI: student guidance https://www.kcl.ac.uk/about/strategy/learning-and-teaching/ai-guidance/student-guidance#section-4
- Generative AI practicals: Using ChatGPT as a dialogic tutor https://youtu.be/iCRe-lZJ9ds?si=XK6WkRMkRwoK1InG
- Meso-level guidance: departments, programmes and modules https://www.kcl.ac.uk/about/strategy/learning-and-teaching/ai-guidance/meso-level