Machine learning development qualifies for R&D tax relief when it seeks an advance in the field of computer science or machine learning, through resolving uncertainty that a competent professional could not readily resolve from published knowledge. Fine-tuning a well documented model on your own data, following the vendor’s playbook, does not usually meet that test. Pushing past what published architectures, training methods or deployment toolchains can demonstrably do often does.
The distinction matters because ML claims sit inside the software category that HMRC scrutinises most closely. This article sets out where the line falls, and what evidence keeps a genuine ML claim on the right side of it. The underlying definition is covered in full in what counts as qualifying R&D.
What does “an advance in the field” mean for machine learning?
It means extending what the field can do, not what your company can do. The DSIT guidelines require an advance in the overall knowledge or capability of a field of science or technology. Applying a published state-of-the-art model to a commercial problem it has not met before is usually an advance for your business only: the field already knew the technique worked.
The question to ask of any ML project is concrete. Could a competent ML engineer, with access to the published literature, model documentation and standard tooling, have said in advance how to hit your targets? If yes, the work was development, not R&D, however valuable the product. If no, and you can say why not, you are probably looking at a qualifying project.
Commercial novelty proves nothing either way. Being first in your market with an ML-driven product is a business fact, not a technological one.
Which kinds of ML work usually reach the bar?
Three recurring territories, described at concept level.
Novel architectures and training methods. Where published approaches cannot meet the accuracy, reliability or latency targets the problem demands, and the team designs or substantially modifies architectures, loss functions or training regimes to get there, the outcome is genuinely uncertain. The published benchmarks define the baseline; the work is the systematic attempt to beat it.
Difficult data regimes. Much published ML performance assumes large, clean, well labelled datasets. Real problems often offer sparse, noisy, heavily imbalanced or restricted data on which established methods degrade unpredictably. Systematic experimentation to achieve reliable performance in such a regime, where the literature cannot say what will work, can qualify. Collecting and labelling data by standard methods, on its own, cannot.
Deployment constraints. Getting a model to run within a hard memory, latency or power budget, on edge hardware or inside a real-time system, can involve genuine uncertainty where documented compression, quantisation and optimisation techniques demonstrably fall short of the target. If the vendor toolchain gets you there by its documented route, it does not.
In each case the test is the same: the uncertainty must be technological, resident in the field, and beyond ready resolution by a competent professional. Difficulty is not enough. Long training runs and expensive compute can be entirely predictable, and predictable work does not qualify.
What ML work does not qualify?
The recurring non-qualifiers in claims we review:
- fine-tuning or prompt-engineering a documented model by established practice, however good the result
- integrating hosted model APIs into a product
- standard MLOps: pipelines, orchestration, monitoring, dashboards, labelling operations
- routine data cleaning and feature engineering using known methods
- projects whose only uncertainty was commercial: whether users would adopt the product, whether the unit economics would work
Some of these activities can sit inside a qualifying project as support for the experimental work. What they cannot do is carry a claim by themselves.
Why is the baseline harder to establish in ML?
Because the field moves quickly, and the claim must be judged against the field as it stood during the accounting period, not as it stands when the claim is written. A capability that was genuinely open in the year the work was done may be a solved problem eighteen months later. Write the baseline down at the time, with references to what was published and what the team tried first, and the later claim inherits that credibility. Reconstructing it afterwards invites hindsight bias, and HMRC’s reviewers are alert to baselines that quietly assume today’s knowledge.
This is also where the competent professional test does its work. HMRC expects the claim anchored in the judgement of a named senior technical person, and the Additional Information Form requires the senior internal R&D contact to be named on every claim. When we prepare an ML claim, that person’s account of what the field could and could not do is the spine of the narrative.
What experiment evidence does HMRC expect?
The ordinary artefacts of real ML research, kept as you go: experiment tracking records, training logs, ablation results, evaluation runs against the baseline, dataset version histories, and notes on why approaches were abandoned. Failed experiments are not an embarrassment in a claim; they are often the clearest proof the uncertainty was real.
Teams running a disciplined experiment tracker already hold most of this. The gap is usually the connective tissue: a record of what question each experiment was asking and what the result changed. A short running log alongside the tracker turns raw runs into evidence.
On the cost side, ML claims draw on staff time (apportioned to the qualifying work), and on software, data and cloud computing costs, which are claimable categories for current periods. The detail is in which costs qualify for R&D tax relief.
How closely does HMRC look at AI and ML claims?
Closely, and claimants should plan on that basis. HMRC checks roughly one in six R&D claims and has more than 500 staff on R&D compliance, and software claims have long attracted particular attention because weak ones are common. HMRC publishes dedicated guidance on software projects, reinforcing the distinction between commercial and technological progress, and an ML claim written in product language rather than in terms of baseline, uncertainty and experiment tends to read as exactly what it is.
None of that should deter a company doing genuine ML research. It changes how the claim is prepared, not whether it is worth making. Our page on HMRC R&D enquiries covers what a check involves and how defensible preparation changes the outcome.
For a loss-making, R&D-intensive AI company, the stakes are worth the discipline: ERIS pays up to 26.97p per £1 of qualifying spend in cash.
Where to go next
The sector picture, including HMRC’s posture on AI claims and how we approach them, is on our AI and robotics sector page. If your uncertainty lives in hardware as much as models, see robotics prototyping and technological uncertainty.
If you are unsure which of your ML projects would survive the test, talk it through with a chartered adviser. A short technical conversation usually settles it.
Written by Matthew Jones ACA CTA. Last reviewed July 2026.
Sources
- DSIT Guidelines: meaning of R&D for tax purposes — the tests of advance in a field, technological uncertainty and the competent professional.
- CIRD81960: Guidelines applied to software — the advance must lie in the underlying technology, not the product.
- HMRC’s approach to R&D tax reliefs 2023–24 — HMRC’s compliance focus, with 17% of claims checked and 500-plus staff on R&D compliance.
This article describes the rules as they stood at the review date above. The rules change: for the current position, start with our guides or talk to us.