arXiv cs.AI→ original

SageMath Improved LLM-Agents in Mathematics by 9.7pp on Average

Researchers integrated SageMath (computer algebra system) into LLM-agents and tested on frontier models like GPT-5.5. All tested models improved performance — by 9.7pp on average. GPT-5.5 achieves 75.2% solvability rate on mathematical problems, Qwen 3.7-Max obtains the maximum gain (+27.8pp). Conclusion: algebra systems in LLM-agents are a promising direction for automated discovery of new mathematical hypotheses.

AI-processed from arXiv cs.AI; edited by Hamidun News
SageMath Improved LLM-Agents in Mathematics by 9.7pp on Average
Source: arXiv cs.AI. Collage: Hamidun News.
◐ Listen to article

Researchers proposed a methodology for integrating SageMath (a computer algebra system) with LLM-agents to solve research mathematics problems. Results showed universal performance improvement of +9.7pp on average, with maximum gains up to 27.8pp for model Qwen 3.7-Max.

How the Algebra System Was Embedded in the LLM

SageMath — free software for mathematical computation, including symbolic algebra, number theory, and visualization. Researchers embedded it in a ReAct-style agent — an architecture that allows LLM to reason step-by-step, calling tools and analyzing results.

Integration includes:

  • ReAct architecture for step-by-step logical reasoning
  • SageMath for verifying and executing algebraic operations
  • Context7 for access to current SageMath documentation
  • RealMath benchmark with research mathematics problems

This combination allows the LLM not just to "guess" answers, but to verify intermediate results and guarantee correctness of computations.

What Results Models Showed

On testing frontier models, the system showed steady performance gains. GPT-5.5 leads in absolute performance: 75.2% solvability and lowest token consumption among all tool-equipped configurations.

Model Qwen 3.7-Max received the maximum relative gain: +27.8pp thanks to SageMath integration. The range of improvements across all tested models ranged from 1.5pp to 27.8pp.

Researchers also proposed an improvement to the RealMath benchmark itself — added multi-step verification and validation to ensure extracted problems were higher quality and more reliable.

Key figures:

  • Average improvement: +9.7pp
  • Range of gains: 1.5pp – 27.8pp
  • GPT-5.5 reaches 75.2% solvability
  • Qwen 3.7-Max receives maximum gain from SageMath

Why Mathematicians Need This

Mathematicians and AI researchers are seeking ways to automate the research process — that very "computational cycle" where a scientist formulates a hypothesis, verifies calculations, analyzes results, and moves to the next idea.

CAS-agents can substantially speed up this cycle: verify complex algebraic derivations in seconds, find patterns in numerical sequences, help formulate and test hypotheses, avoid computational errors.

The authors see in this a path toward automated discovery of new mathematical hypotheses — an area where LLMs previously often erred or required much human help.

What This Means

The research demonstrates a fundamental truth: the right tool choice can qualitatively change an LLM's ability to logical reasoning. SageMath — not the first algebra system, but this is the first major study of its integration into frontier LLM-models at performance scale. The next step — expand the set of CAS-systems (Mathematica, Maple) and optimize integration for other problem types in physics and engineering.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…