Jul29
There is one question I hear almost every time someone demonstrates an AI application. It doesn’t matter whether the application is generating reports, analysing contracts, coaching employees or answering customer questions. Within a few minutes, someone asks, “Which LLM are you using?”
It is a perfectly reasonable question. Large Language Models have transformed artificial intelligence and made AI accessible to millions of people. They have become the public face of modern AI, so it is natural that many people associate the quality of an AI application with the model behind it.
The problem is that the discussion often ends there. It creates the impression that if two organisations use the same LLM, they should achieve similar results. According to McKinsey’s Superagency in the Workplace report, only 1% of organisations believe they have reached AI maturity (Reference: McKinsey – Superagency in the Workplace (2025), despite widespread investment in artificial intelligence. If selecting the right language model were the biggest challenge, that number should be much higher. Clearly, something else separates an impressive AI demonstration from an AI system that consistently delivers business value.
ChatGPT deserves enormous credit for changing how the world thinks about artificial intelligence. It introduced AI to millions of people who had never interacted with machine learning before and demonstrated capabilities that previously felt like science fiction. For many organisations, it became the catalyst for exploring AI projects seriously.
However, artificial intelligence did not begin with ChatGPT, nor did it begin with prompts. Long before Large Language Models became mainstream, organisations were already using AI to detect fraud, recommend products, optimise supply chains, recognise speech, analyse medical images and predict equipment failures. These systems rarely attracted public attention, but they quietly solved real business problems every day.
Generative AI has expanded what machines can do, but it has also unintentionally narrowed how many people think about AI. Today, it is common to reduce an AI application to a prompt and an LLM. That is like judging a modern aircraft solely by the cockpit instruments and displays while ignoring the aerodynamics, engines, flight controls and navigation systems that actually keep it in the air.
While building DealCraft, an AI platform for evaluating enterprise sales conversations, I deliberately started with the Large Language Model. Like many AI engineers, I wanted to understand how much improvement could realistically come from changing the model and refining the prompts before introducing additional layers of intelligence into the system.
For several weeks, I experimented with different language models, rewrote prompts repeatedly and compared how each variation influenced the quality of the evaluation. Better prompts produced better responses, and changing the model occasionally improved consistency. But after a while, each change produced only small improvements. The LLM was doing exactly what it was designed to do, so I realised the next improvement would have to come from somewhere else.
The real breakthrough came when I shifted my attention away from the language model and back to the problem itself. I stopped asking how to make the model generate a better answer and started asking how an experienced enterprise sales leader actually evaluates a customer conversation. That change in perspective completely changed the direction of the project.
Instead of prompt engineering, I found myself working on questions that had nothing to do with the LLM. What evidence should the AI extract from a conversation? Which behaviours genuinely indicate a high-quality enterprise sales discussion? How much should customer discovery influence the final evaluation compared to stakeholder mapping, business value, technical understanding or commercial qualification?
Those weren’t prompt engineering problems. They weren’t language model problems. They were mathematical modelling problems.
The LLM could identify patterns throughout a conversation. It could recognise that business value had been discussed, that stakeholders had been identified or that technical concerns had been raised. Those were useful observations, but they were still only observations.
The difficult part was deciding what those observations actually meant. Should strong customer discovery outweigh excellent product knowledge? If commercial urgency was missing, should the score reduce slightly or significantly? If only half the required evidence existed, how much confidence should the system have in its recommendation?
The language model could not answer those questions because they were never language problems in the first place. They required mathematical models that represented how experienced sales professionals assess opportunities in the real world. Every adjustment to the weighting model, confidence calculation and scoring logic improved the quality of the evaluation far more than another round of prompt engineering.
The LLM produced observations. Mathematics transformed those observations into decisions.
The intelligence comes from mathematics.
Mathematics Turns Data into Decisions
People often describe data as the fuel for AI. I agree. But the engine has to be designed for that fuel.
The same principle applies to artificial intelligence. Raw data is simply a collection of observations until mathematics gives those observations meaning. An AI system might detect twenty different signals in a sales conversation, but someone still has to determine which signals matter, how much they matter and how they should influence the final outcome.
This is where domain expertise becomes software. Years of practical experience are translated into weighting models, scoring functions, confidence calculations and decision rules that a machine can execute consistently. The language model contributes reasoning and pattern recognition, but the mathematical model determines how that reasoning is converted into business decisions that people can trust.
The same principle applies well beyond sales. Credit scoring, fraud detection, recommendation engines, predictive maintenance and medical decision support all rely on mathematical models that convert observations into decisions. Large Language Models have added an extraordinary new capability, but they have not replaced the mathematics that sits underneath intelligent systems.
AI Engineering Is About Turning Expertise into Mathematics
The more AI projects I work on, the more I believe that AI engineering is not primarily about integrating language models. It is about converting human expertise into mathematical models that machines can execute consistently and at scale. The language model becomes one component within a much larger decision system rather than the system itself.
Whether the application evaluates sales conversations, recommends financial products, detects fraud or prioritises customer support cases, the engineering challenge is remarkably similar. Experts naturally understand which factors matter and how those factors influence a decision. AI engineering requires translating that expertise into formulas, weights, probabilities, confidence models and validation rules that software can apply thousands of times without inconsistency.
This part of AI development is rarely visible during product demonstrations because users only see the final answer. They rarely see the months spent understanding the problem, modelling expert judgement, validating assumptions and refining the mathematical relationships that determine how the AI reaches its conclusion. Yet this hidden work often creates far more business value than changing from one language model to another.
We Need to Ask Better Questions
There is nothing wrong with asking which LLM powers an AI application. It is an interesting technical question, and in some situations the answer genuinely matters. It simply should not be the first question or the only question.
A more useful conversation starts by asking how the team decided what the AI should optimise. How were the weighting models developed? What evidence supports the confidence calculations? How were those assumptions validated against real outcomes? How does the system continue learning as new information becomes available?
Those answers reveal far more about the maturity of an AI system than the name of the underlying language model. They also explain why two organisations using exactly the same LLM can produce dramatically different business results.
Final Thoughts
Large Language Models have fundamentally changed artificial intelligence, and there is no doubt they will continue shaping the future of software. They deserve the attention they receive because they have made AI accessible in ways that few technologies ever have. However, reducing every AI application to “just another LLM wrapper” overlooks where much of the real innovation actually happens.
The next time someone demonstrates an AI application, ask which model it uses if you are curious. Then ask a second question that matters even more: Where does the intelligence come from? If the answer begins and ends with the LLM, you are probably looking at a demonstration. If the answer includes mathematical models, data, engineering and domain expertise working together, you are looking at an AI system designed to solve real business problems.
Large Language Models made AI accessible. Mathematics is what makes AI useful.
By Sajeed Ahmed
Keywords: AI, Generative AI, Sales
The Intelligence Comes from Mathematics: Why AI Engineering Is More Than an LLM Wrapper
The Synergy of Agency and Control: How NVIDIA's AI Super-Factory Needs the Singularity Equation
Stop Handholding Your Future Leaders!
The Call - Looking Back After 50 Years
How to succeed with AI adoption?