What Is Actually Running: Autonomous Haulage at Scale
Start with the part of mining AI that is not in dispute. Autonomous haulage is no longer a pilot programme, a proof of concept, or a slide in a technology roadmap. It is production infrastructure moving material at a scale that is hard to argue with, and it has been for long enough that the cumulative tonnage figures have gotten genuinely large.
On 21 April 2026, Komatsu commissioned its 1,000th ultra-class autonomous haul truck— a 930E-5AT with a 290-tonne payload — at Barrick's Nevada Gold Mines. Across the installed base, customers running Komatsu's autonomous haulage system have moved 11.5 billion metric tonnes of material. That is the single most useful number in this entire category, because it is a count of physical work performed, published by the company that performed it, and it does not require you to believe anyone's estimate of a counterfactual.
Caterpillar is on a similar trajectory from a different starting point. It reported 690 autonomous trucks in operation at the end of 2024 and has publicly stated a target of more than 2,000 by 2030. Its named deployments are concentrated in oil sands and aggregates: Suncor's Base Plant, Fort Hills, Imperial Oil's Kearl operation, where 81 Cat 797s were converted to fully autonomous operation in 2023, and a Luck Stone quarry that passed one million tonnes hauled autonomously in July 2025. In March 2026, Fortescue extended its Caterpillar autonomy agreement across three Western Australian sites.
Rail is the under-discussed half of the story. Rio Tinto's AutoHaul system operates at Grade of Automation 4 — no driver on board — with up to 53 trains running at once, each around 240 wagons, roughly 2.5 kilometres long, carrying approximately 28,000 tonnes. The Pilbara network moved 326.2 Mt in 2025. A heavy-haul railway running driverless across a 2,000-kilometre network is a more complete autonomy story than anything happening in a pit, and it gets a fraction of the attention.
| System | Deployment | Cumulative or annual measure | Source |
|---|---|---|---|
| Komatsu FrontRunner — ultra-class autonomous haul trucks | 1,000th unit commissioned 21 April 2026 at Barrick's Nevada Gold Mines; a 930E-5AT with a 290-tonne payload | 11.5 billion metric tonnes of material moved autonomously to date | Komatsu newsroom, April 2026 |
| Caterpillar Command for hauling — fleet size | 690 autonomous trucks at the end of 2024 | Stated target of more than 2,000 autonomous trucks by 2030 | Caterpillar Investor Day 2025, as reported by International Mining |
| Caterpillar — named deployments | Suncor Base Plant, Fort Hills, Imperial Oil Kearl (81 Cat 797s fully autonomous, 2023), and a Luck Stone quarry | Luck Stone passed one million tonnes hauled autonomously in July 2025 | Caterpillar / operator disclosures |
| Rio Tinto AutoHaul — heavy-haul rail, GoA4 | Up to 53 driverless trains, 240 wagons, around 2.5 km long, roughly 28,000 tonnes per train | 326.2 Mt moved across the Pilbara network in 2025 | Rio Tinto 2025 Annual Report |
| Fortescue — Caterpillar autonomy agreement | Extended across three Western Australian sites in March 2026 | Expansion of an existing programme, not a first deployment | Fortescue / Caterpillar announcement, March 2026 |
Verified autonomous haulage deployments in mining as of August 2026. Every figure traces to an OEM release, an operator disclosure, or an annual report.
This article is written for the person who has to decide what to fund next. It is not an argument against mining AI. It is an argument for knowing which parts of it have evidence behind them, because the gap between the two halves of this category is wider than in almost any industry we work in.
The Productivity Numbers That Do Not Survive Sourcing
Here is the finding that shaped this article, and it is not one we expected. When we went looking for the productivity numbers that everyone in mining technology quotes — the 15 to 20 percent improvement over manned operations, the specific figures attributed to individual miners — none of them could be traced to a first-party source.
Not from Rio Tinto. Not from BHP. Not from Fortescue. Not from Caterpillar or Komatsu. Not in an annual report, an investor presentation, a technical paper, or a press release. Every trail ended at a consultancy blog post, an AI-generated listicle, or a number from around 2019 that has been reprinted so many times it has acquired the texture of a fact.
This is not a claim that autonomous haulage does not improve productivity. It almost certainly does — the operational logic of removing shift changes, crib breaks and human variance from a haul cycle is sound, and companies do not commission a thousandth truck for fun. It is a claim about what you can defensibly write in a capital request, put in front of a board, or use to hold a vendor accountable. And the honest answer is: not those numbers.
| Circulating claim | What it actually resolves to | What you can say instead |
|---|---|---|
| Autonomous haulage delivers 15–20% productivity versus manned operations | Consultancy blog posts and AI-generated listicles. No first-party source from any operator or OEM. | Komatsu has commissioned 1,000 ultra-class autonomous trucks and its customers have moved 11.5 billion tonnes autonomously. State the scale. |
| Fortescue reports a 30% productivity gain from autonomy | Secondary reporting with no traceable Fortescue disclosure behind it | Fortescue extended its Caterpillar autonomy agreement across three Western Australian sites in March 2026. State the commitment. |
| Caterpillar reports up to 15% productivity improvement | Recycled marketing copy; not located in any current Caterpillar filing or release | Caterpillar had 690 autonomous trucks at end-2024 and has stated a target of more than 2,000 by 2030. State the trajectory. |
| Rio Tinto operates 400+ autonomous trucks | Figures of this shape date to roughly 2019 and are repeated without a date | AutoHaul runs up to 53 driverless trains and the Pilbara network moved 326.2 Mt in 2025. Use the number that has a report behind it. |
| Two-digit downtime reductions from predictive maintenance on mobile fleet | Vendor and consultancy content; no named operation, no measurement window, no baseline | No measured public result for mining predictive maintenance could be found from a primary source. Say that, and run your own baseline. |
| A named miner saved $5.5M predicting haul-truck maintenance | Attributed to two entirely different companies by different sources — the attribution is actively contradictory | Say nothing. A case study that cannot decide whose case it is has no evidentiary value. |
| Recovery improvements from machine-learning ore-grade control | No named mine, no published methodology, no confidence interval | Describe the mechanism — tighter reconciliation between block model and mill feed — and measure it yourself. |
| Explosives and blasting-cost reductions from machine learning | Content-farm material, with figures that are not plausible as general results | A published random-forest model predicted fragmentation at R-squared 99.2% on test data. Cite predictive accuracy, not cost savings. |
Mining AI claims audit — each circulating figure traced to its furthest verifiable origin, August 2026.
There is a specific pathology worth naming, because it will show up in vendor decks you receive this quarter. A widely repeated case study describes 5.5 million dollars saved through haul-truck maintenance prediction. Different sources attribute that same figure to two entirely different companies. It is not that one attribution is right and the other is a typo — it is that the number has travelled far enough from its origin that nobody carrying it knows whose it was. A figure in that condition has no evidentiary value at all, and its appearance in a deck is a clean test of whether anyone at the vendor checked anything.
If a number in a vendor deck cannot be traced to a filing, an operator disclosure, a peer-reviewed paper, or an explicit our-estimate label, it is not a number. It is a sentence shaped like one.
— Frenchy Digital sourcing principle
The practical consequence for a technology manager is that you have to build your own baseline. There is no industry benchmark to borrow, because the benchmarks in circulation are not measurements. That is more work, but it has a compensating advantage: a baseline you measured is a baseline you can defend when someone asks whether the project delivered.
How to write the capital request without an industry number
Replace the borrowed percentage with three things you can source. First, the deployment scale — a thousand ultra-class autonomous trucks commissioned and 11.5 billion tonnes moved is evidence of technical maturity that no percentage adds to. Second, the safety exposure — powered haulage caused 13 of 33 US mining fatalities in 2025, from a federal count. Third, your own measured baseline for the specific process you are changing.
Then state the expected benefit as a hypothesis with a measurement plan attached rather than as a projection with a decimal place. Reviewers who have seen a few technology business cases find that more credible, not less, and it protects you at the point where someone asks, eighteen months later, where the original number came from.
Safety Is the Honest Argument for Autonomy
There is a well-evidenced argument for taking the operator out of the cab, and it is not about productivity. It is about the fact that powered haulage keeps killing miners.
MSHA recorded 33 mining fatalities in the United States in 2025, up 27 percent from 26 in 2024. Twenty-five occurred at metal and nonmetal operations and eight at coal. Powered haulage was the leading cause, accounting for 13 deaths — more than any other category, and more than twice the six attributed to machinery.
| Measure | 2025 | Context | Source |
|---|---|---|---|
| Total US mining fatalities, 2025 | 33 | Up 27% from 26 in 2024. Preliminary as of early January 2026. | MSHA |
| Metal and nonmetal fatalities, 2025 | 25 | Where most autonomous-haulage-eligible operations sit | MSHA |
| Coal fatalities, 2025 | 8 | Separate regulatory regime and a different equipment mix | MSHA |
| Powered haulage fatalities, 2025 | 13 | The single leading cause, and the exact hazard class that removing the operator from the cab addresses | MSHA |
| Machinery fatalities, 2025 | 6 | Second-largest identified category in the same reporting | MSHA |
US mining fatalities in 2025 by category. The 33 total was preliminary as of early January 2026; confirm against MSHA fatality reports before citing it in a formal document.
Note the flag on that number honestly: the 33 figure was preliminary when it was first reported in early January 2026, and MSHA's final accounting can move it. That caveat is worth carrying, because the discipline of stating it is the same discipline that keeps the rest of your numbers clean.
The fatality increase from 26 to 33 also complicates a narrative you will hear. Autonomy has been scaling for years, and 2025 was worse than 2024. That does not mean autonomy failed — the fleets are concentrated at a small number of very large operations, and most of the industry is nowhere near them. But it does mean nobody should be presenting autonomous haulage as an industry-level safety trend. It is a site-level control at the sites that have it.
What this means for how you frame the business case internally
Lead with the hazard. Powered haulage caused 13 of 33 mining fatalities in 2025, and an autonomous system removes the operator from that hazard entirely. That is a statement you can support with a federal data source, and it is the statement that will still be true when someone audits your business case in two years.
Then treat productivity as a hypothesis you will measure rather than a benefit you will claim. Instrument the haul cycle before anything changes: cycle times, queue times, availability, utilisation, tonnes per operating hour, by fleet and by shift. If the improvement is real, your own data will show it and you will own the number. If it is smaller than the brochure suggested, you will have learned that on your own terms.
The MSHA Silica Correction Most Compliance Plans Get Wrong
If you are building or buying anything that touches occupational exposure monitoring, this section will save you a rework cycle. A great deal of published guidance — including guidance written in 2025 and 2026 — still describes the 2024 MSHA silica thresholds as if they were in force. They are not.
The 2024 rule set a permissible exposure limit of 50 micrograms per cubic metre as an eight-hour time-weighted average, with an action level of 25 micrograms per cubic metre. Original compliance dates were 14 April 2025 for coal and 6 April 2026 for metal and nonmetal. Then two things happened. The Eighth Circuit ordered a judicial stay of the compliance deadlines on 11 April 2025. And MSHA subsequently published a final rule on 6 April 2026 indefinitely delaying the conforming amendments pending litigation.
| Element | What it says | Current status | What to do about it |
|---|---|---|---|
| Permissible exposure limit under the 2024 rule | 50 µg/m³ as an 8-hour time-weighted average | Not currently enforceable | Design monitoring so the threshold is configuration, not code |
| Action level under the 2024 rule | 25 µg/m³ | Not currently enforceable | Keep sampling at this resolution anyway — it is your baseline if the rule returns |
| Original compliance dates | Coal 14 April 2025; metal and nonmetal 6 April 2026 | Superseded by the judicial stay | Remove these dates from any compliance calendar that still shows them as live |
| Eighth Circuit judicial stay | Compliance deadlines stayed on 11 April 2025 | In effect | This is the event most published guidance still omits |
| MSHA final rule, 6 April 2026 | Indefinitely delays the conforming amendments pending litigation | In effect | Cite the Federal Register document, not a summary, when briefing your board |
| Standards actually in force | 30 CFR 56.5001, 56.5005, 57.5001, 57.5005 | In force and enforced | These are the thresholds your programme must currently meet |
MSHA respirable crystalline silica — regulatory status as of August 2026. Cite the Federal Register document directly rather than a summary.
The engineering consequence is small and specific: build the threshold as configuration, never as a constant in code or a hard-coded rule in a compliance module. The regulatory number has now changed twice in two years and is subject to active litigation. A monitoring system where updating the action level requires a release is a system that will be wrong for whatever period your release cycle takes.
The operational consequence is that you should not stand down your sampling programme. A judicial stay is a pause in enforcement, not a finding about exposure, and the litigation can resolve in either direction. Keep collecting at the finer resolution. If the 2024 thresholds return, you already have the historical baseline that shows where each work area stood, which is worth considerably more than the sampling cost. If they do not return, you have a better dataset than the operator next door.
One honest gap: we could not verify the current state of MSHA rulemaking specific to AI, automation, proximity detection or powered haulage. If your plan depends on an anticipated rule in that space, check MSHA's Unified Agenda directly rather than relying on secondary summaries — including this one.
Mining AI Beyond Haulage: Mechanism, Not Savings
Outside autonomous haulage, the public evidence base for mining AI is thin enough that the honest framing changes. In haulage you can point at a thousand trucks. In drill-and-blast, predictive maintenance, ore-grade estimation and ventilation, you can point at mechanisms and at academic work — and you should stop there, because that is where the sourcing stops.
Two areas have genuinely citable peer-reviewed results. Both are worth understanding in detail, because they show what a real result looks like and, by contrast, what the rest of the category is missing.
Drill and blast. A published random-forest model for blast fragmentation prediction achieved an R-squared of 99.5 percent on training data and 99.2 percent on testing data. Field fragmentation across the study ranged from 0.15 to 0.82 metres, with 0.45 metres identified as optimal for a 12 cubic metre shovel bucket. Read what that result actually is: a demonstration that fragmentation can be predicted accurately from blast design and rock mass parameters. It is a well-posed supervised learning problem with a measurable ground truth, which is exactly why the evidence exists here and not in the areas where the outcome is diffuse.
Tailings deformation. Work combining Sentinel-1 InSAR with LSTM models enables millimetre-scale deformation detection over tailings facilities. The more interesting capability is interpretive rather than metric: the approach can separate ordinary consolidation settlement from shear deformation, which is the pattern that matters as an instability precursor. Distinguishing benign settlement from the deformation signature that precedes failure is the actual analytical problem, and a method that addresses it is a real contribution.
| Application area | What is published and citable | What could not be verified | How to read it |
|---|---|---|---|
| Drill and blast optimisation | A random-forest model achieved R-squared of 99.5% on training and 99.2% on testing data. Field fragmentation ranged 0.15–0.82 m, with 0.45 m identified as optimal for a 12 m³ shovel bucket. | Any percentage reduction in explosive consumption or blasting cost. The circulating figures have no primary source. | Fragmentation prediction is a well-posed supervised learning problem with a measurable ground truth. That is why the evidence exists here and not elsewhere. |
| Tailings deformation monitoring | Sentinel-1 InSAR combined with LSTM models enables millimetre-scale deformation detection and can separate consolidation settlement from shear deformation as an instability precursor. | Any claim of failure prediction, or a stated lead time before an event | Satellite interferometry has a public data source and a physical observable. Deployment belongs under the Engineer of Record, as an added signal. |
| Predictive maintenance on mobile fleet | Nothing measured and public from a primary source | Any downtime reduction figure; the widely repeated $5.5M case study is attributed to two different companies by different sources | Sensor data exists but results are commercially sensitive. Run your own baseline before and after — there is no external benchmark to borrow. |
| Ore grade estimation and reconciliation | Nothing measured and public from a primary source | Any recovery improvement figure — none carries a named mine or a methodology | The mechanism is credible — tighter reconciliation between block model and mill feed — but nobody has published the number. |
| Ventilation on demand | No results figures found at all | Any energy or cost saving figure | Underground-specific, capital-intensive, and the published record is essentially empty. Treat vendor claims here as entirely unverified. |
Mining AI application areas mapped against the public evidence base, August 2026. Absence of evidence is stated rather than filled in.
For a technology manager, this reframes how you evaluate a proposal in these areas. You cannot compare a vendor's claim to a benchmark, because no credible benchmark exists. What you can do is require the vendor to describe the mechanism, name the observable it predicts, state the ground truth it is validated against, and agree in advance on how success will be measured at your site. A vendor with a real system will engage with that conversation. A vendor selling a percentage will not.
Tailings: GISTM Conformance and What InSAR Actually Detects
Tailings is the one area where the governance framework is more mature than the technology conversation, which makes it unusually easy to scope correctly. The Global Industry Standard on Tailings Management defines six topic areas, 15 principles and 77 auditable requirements covering the full facility lifecycle — siting, design, construction, operation, closure and post-closure — with governance and information management running across all of it.
The timeline has now fully arrived. ICMM members committed to conformance for extreme and very high consequence facilities by August 2023 and for all other applicable facilities by 5 August 2025. That deadline has passed and disclosure is live. The Global Tailings Management Institute launched in January 2025 with UNEP and the Principles for Responsible Investment to run independent auditing and certification, which moves the standard from a commitment into an audited regime.
This changes what a monitoring system needs to produce. Under an auditable standard, the output is not a dashboard — it is evidence. Every alert needs a provenance chain: what data source, at what time, processed how, reviewed by whom, escalated to whom, closed on what basis. If your monitoring stack cannot reconstruct that chain for a specific date six months later, it will not survive an audit regardless of how good the detection is.
- Map every output to a requirement: The 77 auditable requirements are your specification. An alerting feature that does not correspond to a requirement is a feature you will maintain and never be asked about; a requirement with no corresponding output is a finding waiting to happen.
- Treat InSAR as a supplementary signal, not a replacement: Satellite interferometry adds coverage between instrument locations and a temporal record that survives staff turnover. It does not replace piezometers, inclinometers or survey monuments, and no published work claims it does.
- Separate detection from interpretation in the architecture: Millimetre-scale detection is a measurement. Distinguishing consolidation settlement from shear deformation is an interpretation, and interpretation belongs to a qualified person with the model as an input.
- The Engineer of Record stays accountable: GISTM assigns accountability to named roles. Any AI layer sits underneath that structure and feeds it. Nothing in an agent architecture transfers engineering accountability, and any vendor implying otherwise has misread the standard.
- Log false positives as carefully as true ones: An alerting system nobody trusts is worse than no alerting system, because it trains people to dismiss. Track alert disposition from day one and tune against it.
- Retain longer than you think you need: Deformation is a slow signal. A two-year retention window on processed results is not enough to characterise a facility that has been operating for twenty.
One gap we will state rather than paper over: we could not verify current 2026 figures on water management in mining, which is the obvious adjacent domain. If a vendor quotes you water-related savings, treat that number the same way as the haulage productivity claims until they show you where it came from.
The Capital Contradiction Shaping 2026 Budgets
Any technology proposal you write this year lands in a specific capital environment, and it is a genuinely strange one. The demand story and the investment story are moving in opposite directions at the same time.
According to the IEA Global Critical Minerals Outlook 2026, critical mineral investment fell 9 percent in 2025 — the first decline after years of growth. Battery-metals capital expenditure fell more than 20 percent, the largest decline in over a decade, with lithium investment cut by roughly 40 percent. Exploration spending fell more than 10 percent, with declines around 45 percent in lithium and nickel. Copper-focused companies were the exception, increasing spending 8 percent year over year.
At the same time, demand in the IEA's Stated Policies Scenario nearly doubles to 2040, with lithium rising more than threefold. And supply concentration reached a record: the average share of the top refined supplier hit 70 percent in 2025, up from 68 percent in 2020, and top producers accounted for over three-quarters of total supply growth between 2023 and 2025. Rare earths were the only mineral where concentration did not rise.
Practically, this means the shape of a fundable proposal in 2026 is different from 2022. Long-horizon platform investments compete against a capital pool that just contracted. Discrete, measurable projects with a named internal owner and a payback the operations manager can describe in one sentence compete much better. It also means a vendor whose pricing model assumes a rising capital market is reading a different market than the one you are working in, and that is worth noticing during evaluation.
Where Administrative Agents Actually Earn Their Keep
Autonomous haulage is an OEM decision made at the fleet level with a nine-figure horizon. That is not the project most technology managers are actually scoping. The project most people are scoping is administrative: the paperwork, the reconciliation, the reporting, and the low-grade coordination work that consumes supervisor hours without appearing on any production metric.
That work is where a language-model agent is genuinely good, for an unglamorous reason. It has a clear input, a documented rule set, and a human who signs at the end. Those three properties are what make automation safe and measurable, and they are exactly what safety-critical mining decisions lack.
| Workflow | What the agent does | Human checkpoint | Injection exposure |
|---|---|---|---|
| Shift handover and daily production summary | Assembles a draft from shift logs, fleet management exports and the historian; flags variances against plan and lists open items carried forward | Supervisor edits and signs before distribution | Low — inputs are internal systems of record |
| Contractor and visitor compliance checking | Reads submitted inductions, certifications and insurance against a documented matrix; flags expiry, missing documents and mismatched scopes | Site access decisions stay with the responsible person | Medium — documents arrive from outside the organisation |
| Work-order triage and duplicate detection | Normalises free-text work-order descriptions, clusters likely duplicates, proposes priority and trade assignment from history | Planner accepts, edits or rejects; nothing auto-dispatches | Low |
| Permit and inspection paperwork assembly | Pre-fills recurring inspection and permit forms from prior submissions and current asset data; highlights fields that changed | Named signer reviews every field before submission | Low |
| Procurement and invoice reconciliation | Matches invoices to purchase orders and goods receipts, flags price and quantity variance, drafts the query email | Finance approves; the agent never releases payment | Medium — supplier documents are untrusted input |
| Regulatory and disclosure drafting | Drafts recurring submissions from source data and prior filings, with every claim linked to its source record | Legal and compliance review; nothing files automatically | Low |
| Training and competency tracking | Reconciles competency records against role requirements and rosters; flags gaps before they become access problems | Training coordinator confirms | Low |
Administrative agent workflows for mine sites, with the human checkpoint that makes each one safe — Frenchy Digital, 2026.
Start with shift handover if you are choosing one. It is the highest-frequency document on most sites, its inputs already exist in systems of record, its failure mode is a supervisor editing a draft rather than an incident, and the time saving is easy to measure because you can time the current process. It also builds the integration substrate — historian reads, fleet system exports, document publishing — that every subsequent workflow reuses.
One more thing worth saying about measurement, because it is the difference between a project that gets renewed and one that gets quietly dropped. Time the current process before you build. Ten supervisors, one week, a stopwatch on the handover document. That number is your baseline, you own it, and it is the only number in this article you will never have to defend the provenance of.
Legacy Integration Is the Binding Constraint
The constraint on mining AI projects is not model capability. It has not been for a while. The constraint is that the systems holding your data were designed to be operated, not integrated, and the write path is where projects actually fail.
Reading is usually solvable. Most fleet management systems, historians and CMMS platforms can produce an export, a replica, or a reporting endpoint, and a determined integrator gets there in weeks. Writing back — under an audit trail, with a rollback, without violating a licence agreement or a change-control process — is a different problem, and it is the one that turns a nine-week estimate into a nine-month project.
| System class | Read access reality | Write access reality | What an agent should write |
|---|---|---|---|
| Fleet management system (Cat MineStar, Komatsu, Wenco, Modular) | Usually available through reporting exports or a data warehouse feed | Closed. Assume none, and design the agent to produce a document or a queued recommendation instead. | Never. Machine dispatch and control is out of scope for an administrative agent. |
| CMMS / EAM (SAP PM, Maximo, Pronto, Mine-specific) | Read APIs or database replicas usually exist | Often exists but is rate-limited, poorly documented, and requires a licence tier nobody budgeted | Draft work orders written to a review queue, created by a human action |
| Process historian (PI, Aveva, Canary) | Well-supported read access; the best data source on most sites | Not applicable — a historian is not a write target | Not applicable |
| Laboratory information system (LIMS) | Usually exportable; sometimes only as scheduled files | Rarely, and validation requirements make writes a regulated change | Not without a formal validation path |
| Mine planning (Deswik, Vulcan, MineSight) | File-based rather than API-based in most environments | File round-trip at best | Read-only. Treat planning packages as a source, never a sink. |
| Document and records management (SharePoint, Enablon, Intelex) | Generally good API coverage | Generally available | Drafts and metadata, with a human publish step |
| ERP finance (SAP, Oracle, Pronto) | Read access via reporting layer | Available but tightly change-controlled | Draft-only, into an approval workflow |
Mining system integration reality by platform class. Confirm the specific licence tier and endpoint availability for your installed versions before any estimate.
The design pattern that survives this constraint is to make the agent produce artifacts rather than mutations wherever possible. A drafted document, a queued recommendation, a populated form, an exception list — these deliver most of the time saving and require none of the write access. Then add write-back selectively, to the one or two systems where it materially changes the workflow, after you have confirmed the endpoint exists and someone has actually authenticated against it.
Ask what the write path is for every target system, in writing, before anyone estimates the model work. A project that assumes a write API exists and discovers in week six that it does not is a project that has already doubled.
— Frenchy Digital scoping principle
Prompt Injection and the Restraint Boundary
Any agent that reads documents you did not author is exposed to prompt injection, and prompt injection is not a solved problem. On a mine site the untrusted inputs are obvious once you look for them: contractor certifications, supplier invoices, tender responses, third-party inspection reports, emailed correspondence, scanned regulatory notices. Each of those is text written outside your organisation that a model may read as instruction rather than as data.
Nobody has a complete defence. What responsible engineering does is reduce blast radius, and the controls that work are architectural rather than prompt-based.
- Retrieved content is data, never instruction: Structurally separate document content from the instruction channel, and never let a retrieved document expand what the agent is permitted to do.
- Allowlist tools per workflow: A contractor-compliance agent has no reason to reach a work-order endpoint. Capability scoping at the tool layer is the control; a system prompt asking the model to behave is not.
- Deny by default any argument the human did not supply: If a parameter did not come from an authenticated human action or a system of record, the tool call fails rather than inferring.
- No write action without a human step: The agent proposes, a named person accepts. For anything with safety, legal or contractual consequence this is not a configuration option.
- Log every tool call with its arguments: When something goes wrong, the question is what the agent did and on whose instruction. That answer has to exist in a log you control.
- Test injection cases in CI: Maintain a corpus of adversarial documents and run it on every prompt, tool or model change. Map the suite to the OWASP Top 10 for LLM Applications so the coverage is legible to a security reviewer.
Cross-reference the OWASP Top 10 for Large Language Model Applications and the NIST AI Risk Management Framework when you write the control set. Neither is mining-specific, and both are more useful than anything mining-specific currently published.
The 90-Day Pilot That Produces a Defensible Number
Everything above leads to one practical conclusion: since there is no credible external benchmark in this category, the pilot is not a technology trial. It is a measurement exercise that happens to involve technology. Design it that way and you finish with a number you own; design it as a demo and you finish with an anecdote.
Ninety days is enough for one workflow at one site if the baseline work starts in week one rather than being retrofitted at the end. The most common way these projects fail is not a model problem — it is that nobody wrote down what the before-state looked like, so the after-state has nothing to be compared against.
| Phase | What happens | The artifact | The common failure |
|---|---|---|---|
| Weeks 1–2 — Baseline | Time the current process with a stopwatch, count the artifacts produced per week, and record who touches each one. Capture error and rework rates in whatever form they currently exist. | A written baseline signed off by the supervisor who owns the workflow | Skipping this. Without it, every later number is an assertion. |
| Weeks 2–4 — Integration proof | Authenticate against every target system and pull real data. Confirm the write path exists at your licence tier, or confirm it does not. | A working read from each system and a written statement of what can and cannot be written | Accepting a vendor's integration claim instead of testing it |
| Weeks 4–8 — Build and shadow run | The agent produces output that nobody acts on. Humans do the work as before and the drafts are compared afterward. | A comparison set showing where the agent agreed, disagreed and failed | Going live before you know the disagreement rate |
| Weeks 8–12 — Supervised live | Output enters the real workflow, always through a human accept-edit-reject step. Every decision is logged with the time taken. | Override rate, edit rate, rejection rate and review time by reviewer | Measuring acceptance without measuring review time — a fast accept is not a review |
| Week 12 — Decision | Compare against the week-1 baseline on the same measure, with the same definition, counted the same way. | A number you own, with a method you can describe in one paragraph | Redefining the metric after seeing the result |
A 90-day measurement-first pilot structure for a single mine-site workflow — Frenchy Digital, 2026.
Two details do most of the work. The first is the shadow run in weeks four to eight. Running the agent alongside the existing process, with nobody acting on its output, is the cheapest way to learn the disagreement rate before it can cause harm — and the disagreement cases are usually more informative than the agreements, because they show you where the source data is inconsistent rather than where the model is weak.
The second is measuring review time rather than acceptance rate. An acceptance rate near 100 percent looks like success and is more often evidence that the review step has become a click. Record the time between presentation and decision as a distribution, and look at the fast tail specifically. A workflow with a healthy median and a large sub-two-second population is a workflow where the human checkpoint has quietly stopped existing.
How to Evaluate a Mining AI Vendor
Given how much unsourced material circulates in this category, vendor evaluation is mostly a sourcing exercise. Ten questions, sent before the demo rather than after the pilot.
| Question | Acceptable answer | Disqualifying answer |
|---|---|---|
| Name one operation where this ran in production, and who we can call | A named site, a named contact, and a willingness to arrange the call | Confidentiality prevents us naming customers — offered for every reference request |
| What was the measured baseline before deployment, and who measured it | A described measurement method, a time window, and who owned the data | A percentage improvement with no stated baseline |
| Show us the primary source for every number in your deck | Filings, operator disclosures, peer-reviewed papers, or an explicit our-estimate label | Links that resolve to consultancy blogs, listicles or the vendor's own earlier marketing |
| What is the write path into our CMMS and fleet system, specifically | Named endpoints, named licence tiers, named limitations, and an admission of what is read-only | Full integration with all major systems |
| What happens when the model is wrong in this workflow | A described failure mode, a human checkpoint, and a rollback | Our accuracy is 99 percent |
| Can we export the full decision log, including tool calls and inputs | Yes, on demand, machine-readable | You can view it in our dashboard |
| Who owns the data and the derived models | You do, in writing, including on termination | Aggregated and anonymised improvement of our platform |
| Does anything you build touch a safety-critical system or interlock | No, and here is the boundary in the architecture | An evasive answer, or an enthusiastic yes |
| How do you version models and notify us of behaviour changes | Pinned versions, advance notice, a changelog, a rollback path | We always run the latest model |
| What is your position on autonomous action without human review | Nothing acts without a named human signer in a regulated or safety-adjacent workflow | Fully autonomous operations as a headline feature |
Frenchy Digital vendor due-diligence question set for mining AI, 2026.
The third question does most of the work. Ask for the primary source behind every number in the deck, and watch what comes back. A vendor with real deployments will send you an operator disclosure, a filing, or a paper, or will tell you plainly that a figure is their own estimate and explain how they derived it. A vendor without them will send you a link to a consultancy blog post, or will explain that the figure is industry consensus. Industry consensus, in this category, means several people repeating the same unsourced number.
Then ask for one artifact rather than a document set: a redacted decision log for a single workflow, covering a single day, including tool calls and inputs. It answers more than any questionnaire, because a vendor who cannot produce it does not have the logging, whatever the security review says.
Red Flags in Mining AI Procurement
Each of these is a specific, checkable failure rather than a general caution. Several of them can be detected from the first deck.
| Red flag | Why it matters |
|---|---|
| A productivity percentage for autonomous haulage with no operator or OEM source | No verifiable first-party figure exists. A vendor quoting one has copied it from content that copied it from somewhere else. |
| The $5.5M predictive-maintenance case study | Different sources attribute the same figure to different companies. Its presence in a deck is a direct test of whether anyone checked. |
| A compliance module that assumes the 2024 silica thresholds are in force | They are stayed and the conforming amendments are indefinitely delayed. A product built on the wrong thresholds was built without reading the docket. |
| Safety claims presented as productivity claims | Safety is the well-evidenced case. A vendor that leads with productivity and hides safety in a footnote has the argument backwards. |
| An agent proposed for ground control, ventilation setpoints or blast approval | These are safety-critical engineering decisions with a named accountable person. There is no architecture that makes an agent an appropriate decision-maker here. |
| Full integration claimed without naming a write endpoint | The write path is where mining AI projects actually fail. A vendor who cannot name it has not attempted it. |
| Percentage savings quoted with no baseline and no measurement window | An improvement claim without a baseline is not a measurement, it is a preference. |
| An AI tailings monitoring product positioned as replacing instrumentation | GISTM governance assigns accountability to named roles. A monitoring layer supplements the Engineer of Record; it does not substitute for one. |
| No mention of prompt injection when the agent reads external documents | Contractor certificates, supplier invoices and tender documents are untrusted input. A vendor unaware of the exposure has not designed for it. |
| Pricing that assumes a rising capital market | Critical mineral investment fell 9 percent in 2025. A business case that assumes budgets are expanding is reading a different market than yours. |
The Frenchy Digital red-flag list for mining AI buyers, August 2026.
In a category where the headline statistics do not survive sourcing, the vendor who tells you what they cannot prove is more trustworthy than the vendor whose deck has a number for everything.
— Frenchy Digital buyer's principle
What It Costs to Build This Properly
These are the bands Frenchy Digital uses to scope operations AI engagements in 2026. They assume integration assessment is in scope from the start, because integration is where the schedule risk lives in mining and discovering it late is what makes a project expensive.
| Engagement | Range | Timeline | Typical scope |
|---|---|---|---|
| Discovery + workflow audit | $9k–$22k | 2–4 weeks | System inventory, integration and write-path assessment per target system, workflow shortlist, measurement baseline design |
| Single-workflow agent (shift reporting, contractor compliance, work-order triage) | $28k–$70k | 4–9 weeks | One workflow end to end, human review queue, decision logging, evaluation set, rollout to one site |
| Multi-workflow operations platform with system integration | $70k–$180k | 9–16 weeks | Several workflows, CMMS and historian integration, retrieval layer, role model, evals in CI, supervisor dashboards |
| Enterprise / multi-site / regulated build | $180k–$420k+ | 14–24 weeks | Multi-site rollout, audit logging, human-in-the-loop instrumentation, SOC 2 posture, disaster recovery and documentation package |
Frenchy Digital cost bands for mining and heavy-industry AI engagements, 2026.
Senior-led delivery runs $150 to $225 per hour, and ongoing retainers run $2,500 to $9,500 per month covering model and dependency upgrades, evaluation expansion, incident response, and a quarterly technical review. Every engagement carries a 30-day post-launch warranty, and you receive a written scope with a fixed-price phased proposal within 5 business days of the discovery call.
The budgeting note that surprises people is the same one that applies in every regulated industry: the integration substrate is largely a fixed cost, paid once, reused by everything after. The first agent pays for the historian connection, the document pipeline, the identity model, the review queue and the logging. The fourth agent inherits all of it. Sites that sequence their automation get materially better economics than sites that run four disconnected pilots in parallel — which, in a year when critical mineral investment fell 9 percent, is not a small consideration.
Limitations and Honest Failure Modes
This article has argued that the evidence base for mining AI is uneven. That argument applies to what we build as much as to what anyone sells, so here is the honest list.
- There is no benchmark to compare against: Because the circulating productivity and savings figures do not resolve to primary sources, you cannot benchmark a proposal against the industry. Every business case has to be built on a baseline you measured yourself, which is more work and the only defensible option available.
- Autonomous haulage is not the project most sites can run: It is concentrated in very large iron ore, oil sands and gold operations. If your site does not look like those, the relevant question is administrative automation, not fleet autonomy, and conflating the two produces a business case that does not survive contact with a CFO.
- The write path may not exist: Discovery sometimes concludes that the target system cannot be written to at your licence tier, or that the change-control burden exceeds the benefit. We would rather tell you that in week three than build around it for four months. It is the most common reason a scoped workflow gets swapped for a different one.
- Prompt injection is unsolved: Any agent reading contractor documents, supplier invoices or tender responses is exposed. We reduce blast radius through capability scoping, deny-by-default arguments and human approval gates. We do not claim to eliminate the risk, and no one who does should be believed.
- Document quality sets the ceiling: An agent that reads shift logs is bounded by how well shift logs are written. Sites with inconsistent free-text logging get less value than sites with structured entry, and no model fixes an input problem. Sometimes the highest-value output of a discovery is a form redesign.
- The regulatory ground moves: The silica thresholds have changed twice in two years and are in active litigation. Anything encoding a regulatory number needs that number as configuration and a named owner responsible for checking it, or it will silently go stale.
- Adoption is the usual failure, not accuracy: The workflows that stick are the ones a supervisor already wanted help with. The ones that fail are the ones a head-office programme imposed on a crew that was not consulted. Pick the workflow with the frontline champion, even when it is not the one with the best theoretical return.
- We are not a mining engineering firm: Frenchy Digital builds software. Ground control, ventilation design, blast engineering, tailings governance and MSHA compliance decisions belong to qualified professionals, and our systems feed their judgement rather than substituting for it.
None of this argues against building. It argues for building the measurement alongside the system, choosing one workflow with a real baseline and a named owner, and being honest inside your own organisation about which claims in this category have evidence behind them and which do not. In mining specifically, that honesty is worth more than usual — because the alternative is a business case resting on a number nobody can find the source of.
Scoping an AI Build for a Mine Site?
Book a free 60-minute discovery call with Frenchy Digital — a senior-led Black-owned LA agency. You leave with an integration assessment naming the write path for every target system, a measurement baseline design, and a fixed-price phased proposal within 5 business days. Call +1 (424) 272-5601.
Scoping an AI Build for a Mine Site?
Book a free 60-minute discovery call. You leave with an integration assessment that names the write path for every target system, and a fixed-price phased proposal within 5 business days.
1517 S Bentley Ave Unit 204, Los Angeles CA 90025
Frequently Asked Questions
Sources & References
- 1Komatsu — First OEM to Commission 1,000 Ultra-Class Autonomous Haul Trucks (April 2026)↗
- 2International Mining — Caterpillar targets over 2,000 autonomous mining trucks by 2030↗
- 3Rio Tinto — 2025 Annual Report (SEC Form 6-K exhibit)↗
- 4MSHA — Fatality Reports Search↗
- 5Pit & Quarry — MSHA: Industry finishes 2025 with 33 miner fatalities↗
- 6Federal Register — Respirable Crystalline Silica: Delay of Conforming Amendments (April 6, 2026)↗
- 7eCFR — 30 CFR Part 56, Subpart D (Surface Metal and Nonmetal Air Quality)↗
- 8eCFR — 30 CFR Part 57, Subpart D (Underground Metal and Nonmetal Air Quality)↗
- 9IEA — Global Critical Minerals Outlook 2026↗
- 10IEA — Supply concentration, export restrictions and declining investment put critical mineral security at risk↗
- 11ICMM — Global Industry Standard on Tailings Management↗
- 12Global Tailings Review — The Global Industry Standard on Tailings Management↗
- 13Global Tailings Management Institute — About the GTMI↗
- 14UNEP — Global Industry Standard on Tailings Management↗
- 15Natural Resources Research — Random forest prediction of blast fragmentation (Springer)↗
- 16Ain Shams Engineering Journal — Sentinel-1 InSAR and LSTM for tailings deformation monitoring (ScienceDirect)↗
- 17OWASP — Top 10 for Large Language Model Applications↗
- 18NIST — AI Risk Management Framework↗

