
Benchmarks are crucial for assessing the capabilities of Large Language Models (LLMs) in real-world tasks, but developing and updating these benchmarks is both costly and time-consuming. This paper explores an innovative approach using Agentic AI, where LLMs themselves generate and evaluate practical exams for various occupations, including those in Finance & Business Operations, Management, and Computer & Mathematics.
The Agentic AI approach involves distinguishing between the materials needed for tasks (such as text, data, and images) and the tools required to solve them (such as function calling or web search). Focusing on text-only tasks that do not require tool use, the study found that only 7% (149 tasks) of these occupations are currently testable using this method. This limitation highlights the need for further development in benchmarking methodologies that can accommodate a broader range of tasks.
Results and Findings
Thirteen different models, including variants of GPT, Claude, and Gemini, were deployed to complete these synthetic exams. The results revealed that even on basic tasks, current LLMs face significant challenges. Leading models achieved median scores ranging from 65% to 79%, with particularly weak performance in data manipulation and financial calculations. This indicates that while LLMs have made progress, there is still a substantial gap in their ability to handle complex, real-world tasks accurately.
The study also noted a rapid improvement in LLM performance over time. Models released in 2024 averaged a score of 40.5%, while those released in 2025 reached an average of 66%, representing a 26 percentage point gain in just one year. This rapid progression suggests that LLMs are becoming more capable at a faster rate, although considerable work remains in validating these benchmarks and extending them to include tool use.
Conclusion and Future Work
The Agentic AI approach offers a promising solution for creating cost-effective, scalable, and updatable benchmarks for measuring AI workplace capabilities. By extending the “LLM-as-a-judge” paradigm to occupational task assessment, this method can provide a more dynamic and adaptable way to evaluate LLM performance in various professional settings. However, further research is needed to validate these benchmarks and to develop methods that can assess tasks requiring tool use, thereby expanding the scope of testable occupations beyond the current text-only limitations.
The development of LLM-generated benchmarks holds significant potential for the future of AI evaluation. As LLMs continue to evolve, the ability to automatically generate and evaluate practical exams can streamline the benchmarking process, making it more efficient and responsive to the rapidly changing landscape of AI capabilities. This approach not only reduces the cost and effort required to develop benchmarks but also ensures that they remain relevant and up-to-date with the latest advancements in AI technology.
Summary and Implications
In summary, the Agentic AI approach to benchmarking LLM capabilities represents a significant step forward in the field of AI evaluation. By leveraging LLMs to generate and evaluate practical exams, this method offers a scalable and cost-effective solution for assessing AI performance in real-world tasks. While there are still challenges to overcome, particularly in validating benchmarks and extending them to include tool use, the potential benefits of this approach are substantial. As AI continues to advance, the development of more sophisticated and adaptable benchmarking methodologies will be crucial for ensuring that AI systems are evaluated accurately and effectively.
