The Independent Evidence Contradicts the Sales Pitch
Container terminals are unusual among the industries we write about, and the reason is worth stating up front. Most sectors being sold AI have almost no independent evidence about whether the technology delivers — you get vendor case studies, analyst reports commissioned by vendors, and a lot of adjectives. Ports are different. Terminal automation has been running at scale for over a decade, it has attracted attention from a global consultancy, an intergovernmental transport body, peer-reviewed academics, a trucking association with GPS data and the World Bank, and several of those parties have published.
That evidence does not say what the sales deck says. That is the article.
This is written for a terminal operator, a port authority executive or an operations director sitting across the table from a vendor. It is not an argument against automation, and it is certainly not an argument against AI in terminals. It is an argument that the specific claims being made to you have a documented track record of not surviving independent measurement, that you should know which claims those are, and that there is a separate category of work — unglamorous, back-office, document-heavy — where the case is genuinely strong and almost nobody is pitching it to you.
One more framing note. Everything Frenchy Digital builds in this space is administrative and documentation automation under human review. Nothing below proposes an agent that moves a container, releases cargo, allocates a berth or makes a safety-critical maintenance call on its own. The evidence problem in this industry is bad enough without adding autonomy to it.
McKinsey's Practitioner Survey: Expectations Versus Outcomes
Start with the finding that a consultancy with a large logistics practice published about its own clients' industry, because it is the least likely thing to be marketing.
McKinsey's "The Future of Automated Ports"surveyed industry practitioners about what they expected from automation and what they actually got. The gap is the story. Respondents expected automation to cut operating expenses by 25 to 55 percent and to raise productivity by 10 to 35 percent. In the report's own words, "today these expectations generally aren't realized, especially in fully automated projects." Operating expenses did fall — but only by 15 to 35 percent. And then the sentence the industry has never comfortably answered: "Worse, productivity actually falls, by 7 to 15 percent."
| Metric | What practitioners expected | What the survey found | What it means for your business case |
|---|---|---|---|
| Operating expense reduction | 25% to 55% lower | 15% to 35% lower | Real, but landing at the bottom of what the business case assumed |
| Productivity | 10% to 35% higher | 7% to 15% LOWER | The sign flipped. This is the finding the sector has never satisfactorily answered. |
| Return on invested capital | At or above the industry norm of about 8% | Short of that norm by up to one percentage point | A project that clears its hurdle rate on paper and misses it in operation |
| What would justify the capex | — | Opex 25% lower than a conventional terminal, or productivity up 30% while opex fell 10% | Neither threshold was being met by the surveyed projects |
| Moves per crane-hour | — | "Low 20s" automated against "high 30s" conventional | Explicitly a single executive's anecdote in the source, not measured data. Do not use it as a benchmark. |
Expected versus realised outcomes at automated ports, per McKinsey's practitioner survey. Note the limitations stated below before citing any of this.
The capital math in the same report is more damning than the headline. To justify the capital expenditure, McKinsey calculated, "operating expenses of an automated greenfield terminal would have to be 25 percent lower than those of a conventional one or productivity would have to rise by 30 percent while operating expenses fell by 10 percent." Neither condition was being met. Return on invested capital was "falling short by up to one percentage point from the industry norm of about 8 percent." A percentage point does not sound like much until you apply it to a $1.5bn programme over its life.
The same caution applies to the number people most want to quote from it. McKinsey mentions automated terminals running in the "low 20s" of moves per crane-hour against "high 30s" at conventional terminals. That comparison is presented in the source as one executive's anecdote. It is not measured data, it is not a survey result, and reproducing it as a benchmark — which happens constantly — is exactly the laundering problem this article exists to describe.
A seven-year-old opinion survey with no replication is weak evidence about today. It is still stronger evidence than a vendor case study published last month, because it had no commercial reason to say what it said. Rank your sources by incentive before you rank them by recency.
— Frenchy Digital evaluation principle
The OECD/ITF Review — and What Peer Review Actually Found
The second source carries different weight, and it is the one we would put in front of a board. The OECD's International Transport Forum published "Container Port Automation: Impacts and Implications" in October 2021. The ITF is intergovernmental. It sells nothing, it rates nothing, and it has no client relationship with a terminal operator or an equipment manufacturer.
Its first finding reframes the urgency. At the time of the review there were roughly 53 automated container terminals worldwide, representing about 4 percent of global capacity. After more than a decade of automation being described as the future of the sector, ninety-six percent of the world's container capacity is handled some other way. That is not an argument against automating. It is an argument against the specific pressure — everyone is doing this, you are falling behind — that shows up in the first five slides of every deck.
Its second finding is the peer-reviewed one. The ITF cites Ghiara and Tei (2021), whose conclusion is that automation alone "cannot be considered to have a highly significant impact on port terminal performance." That is a null result from academic literature, not a consultancy opinion, and it is compatible with the McKinsey survey rather than contradicting it.
| Finding | What the OECD/ITF review reports | How to read it |
|---|---|---|
| Scale of automation | About 53 automated container terminals worldwide, roughly 4% of global capacity | Automation is a niche, not an industry standard. The pressure to keep up is smaller than the deck implies. |
| Peer-reviewed assessment | Ghiara and Tei (2021): automation alone "cannot be considered to have a highly significant impact on port terminal performance" | The strongest available statement from peer review, and it is a null result |
| APMT-Maasvlakte 2 remote quay cranes, 2019–21 | 25 moves per crane-hour, with "hardly any week" exceeding 30 — lagging other Rotterdam terminals | A flagship automated terminal, measured over two years, underperforming its own port |
| Qingdao | 43 moves per crane-hour reported in January 2020 | Automated terminals span an enormous range. The technology is not the variable that explains it. |
| Labour reduction (Prism Economics, 2019) | 40–50% at TraPac LA, 50% at Patrick's Sydney, up to 85% at Qingdao | [FUNDING UNVERIFIED — likely union-commissioned] Directionally consistent with automation cutting headcount, which nobody disputes |
| Capacity claim (Moody's, 2019) | Semi-automation at Norfolk International Terminal South would add 47% capacity on the same footprint | [RATINGS AGENCY, NOT NEUTRAL] The ITF disputes the basis of the comparison |
| Union-commissioned study | Economic Roundtable, "Someone Else's Ocean" (June 2022): about 572 FTE per year eliminated at two automated terminals, 2020–21 | [FUNDED BY THE ILWU COAST LONGSHORE DIVISION] Read it as advocacy research with a real methodology, not as neutral measurement |
Findings from the OECD/ITF review of container port automation, 2021. Funding and commissioner labels are stated where they are known and flagged where they are not.
The Rotterdam comparison in that table deserves its own paragraph, because it is the most specific operational number in the independent literature. APM Terminals' Maasvlakte 2 facility runs remote-operated quay cranes. Across 2019 to 2021, the ITF reports those cranes averaged 25 moves per crane-hour, with hardly any week exceeding 30 — lagging other terminals in the same port. Over the same period, Qingdao was reported at 43. Both are automated terminals. The eighteen-move spread between them tells you that automation is not the variable explaining terminal productivity, which is precisely what the peer-reviewed null result says.
On labour numbers, and why we label the funding
Two of the figures in that table come from studies with an interested commissioner, and we have labelled both. The Prism Economics headcount reductions — 40 to 50 percent at TraPac Los Angeles, 50 percent at Patrick's Sydney, up to 85 percent at Qingdao — appear in the ITF review, but we could not verify who commissioned the underlying work, and the likeliest answer is a union. The Economic Roundtable's "Someone Else's Ocean," which found roughly 572 full-time equivalents per year eliminated at two automated terminals in 2020–21, states its funder plainly: the ILWU Coast Longshore Division.
Labelling them is not dismissing them. Union-funded research on job losses is often methodologically careful, and the direction of these findings is not seriously contested — automation reduces terminal headcount, which is the entire point of it. But a number produced by a party with a stake in the answer belongs in your analysis with the stake attached.
The same rule applies in the other direction. Moody's 2019 estimate that semi-automation at Norfolk International Terminal South would add 47 percent capacity on the same footprint comes from a ratings agency, not a neutral observer, and the ITF disputes the basis of the comparison. We are not printing it as a fact either.
Vendor Claim Versus Independent Measurement, Side by Side
The cleanest way to show the pattern is to put the two columns next to each other. Every row below pairs something a vendor has published with whatever independent measurement exists for the same thing — and in several rows, the honest answer in the right-hand column is that no independent measurement exists at all.
| Claim area | What the vendor published | What independent measurement shows | The gap |
|---|---|---|---|
| Crane labour time | ABB advertised a 45–55% reduction per crane [VENDOR] | Oliveira and Varela (2017) measured 33%, as cited by the OECD/ITF | The advertised floor was 12 points above the measured result |
| Container-code OCR accuracy | Visy: "market leading accuracy of 97–99.5%" [VENDOR]; Camco: "98% guaranteed" [VENDOR, page not verifiable] | No independent third-party test exists | There is nothing to compare the claim against. That is the finding. |
| Quay crane throughput | Camco/ZPMC at Beibu Gulf Port Qinzhou: "up to 36 moves per hour per crane," 8-second OCR per move [VENDOR PR via trade press] | OECD/ITF on APMT-Maasvlakte 2: 25 moves per crane-hour sustained over 2019–21 | "Up to" is a peak. The independent figure is an average. They are not comparable. |
| Truck turn time after an appointment system | "45% reduction" at GCT Bayonne, widely recycled since 2017 | US EPA states "more than 40%" and attributes it to "GCT estimates" — no baseline, no methodology | The canonical statistic in this category is a self-estimate |
| Terminal capacity from TOS optimisation | Kaleris/Navis N4 named user Port of Helsingborg: "up to 30%" capacity gain in the existing area [VENDOR] | No published metrics; Navis N4's own optimisation modules are marketed without AI claims or figures | A named customer is not a measurement |
| Crane maintenance savings at Rotterdam | "18% lower crane repair costs" and "20% shorter waits" circulate widely | Traceable only to SEO content farms; the Port of Rotterdam's own site frames its 4D digital twin as an aim | REJECTED — we will not print these numbers as fact, and neither should your board deck |
Vendor claims paired with independent measurement where it exists. Vendor-sourced figures are labelled; unverifiable figures are marked as such rather than reproduced as fact.
The first row is the cleanest example in the entire batch, and it is worth stating on its own. ABB advertised a 45 to 55 percent reduction in labour time per crane. Oliveira and Varela measured 33 percent. The measured result did not just miss the advertised range — it landed twelve points below the bottom of it. Notice also that 33 percent is a good result. A third of the labour time on a crane is a real saving that would justify a real investment. The problem is not that the technology did nothing. The problem is that the business case was built on a number nobody had measured.
The last row is where we draw a line. The figures circulating about Rotterdam — 18 percent lower crane repair costs, 20 percent shorter waits — trace back only to content-farm articles with no primary source, while the Port of Rotterdam's own material describes its 4D digital twin as an aim. We will not print those numbers as facts. If they appear in a deck presented to you, ask for the primary source, and watch what happens.
Truck Turn Times: Independent GPS Data Against a Nine-Year-Old Self-Estimate
Truck turn time is the metric drayage carriers care about most, it is the one appointment systems are sold on, and it is the one where independent measurement is available — because the Harbor Trucking Association collects GPS geofence data from its members' trucks rather than relying on terminal reporting.
That data, for July 2017, put the San Pedro Bay average at 90 minutes, with 24 percent of trips exceeding two hours. The terminal-level breakdown is where it gets interesting.
| Terminal / scope | Average turn time | Note |
|---|---|---|
| TraPac, Los Angeles — automated | 114 minutes | Slowest terminal in the dataset |
| Middle Harbor / LBCT, Long Beach — automated | 46 minutes | Well inside the average |
| Matson | 38 minutes | Fastest terminal in the dataset |
| Los Angeles, all terminals | 93 minutes | Port-level average |
| Long Beach, all terminals | 87 minutes | Port-level average |
| San Pedro Bay combined | 90 minutes | 24% of trips exceeded 120 minutes |
Harbor Trucking Association GPS geofence turn-time data, July 2017 — collected independently of terminal reporting.
The automated TraPac Los Angeles terminal measured 114 minutes — the slowest in the dataset — against 38 minutes at Matson and 46 at Middle Harbor and LBCT. Be careful about how much weight this carries: it is one month of one association's data from 2017, terminals differ in cargo mix, chassis availability, gate hours and volume, and a single automated terminal at the bottom of one table is not proof that automation slows trucks down. But it is measurement, collected by a party with no interest in the automation debate, and it points the opposite way from the narrative.
Now the appointment-system statistic, which is the more instructive story.
None of this means appointment systems do not work. It means the number your vendor is quoting is not evidence that they work. And the state of play in the largest US port complex suggests the sector itself is not certain either: Los Angeles and Long Beach still have no unified appointment system across their twelve container terminals. In February 2023 the Harbor Trucking Association and the California Trucking Association formally asked for one interoperable system; the response was supportive and carried no timeline. As of the port's State of the Port address in January 2026, an $8 million California GO-Biz grant is still being used to extend the Port of Los Angeles's Universal Truck Appointment System to Long Beach terminals — in progress, with no published measured turn-time result.
Two clarifications that will save you an embarrassing slide. PortCheck is not an appointment system: per its own description it was founded in 2008 as an online rate-collection service for the Los Angeles/Long Beach Clean Truck Fund, and it collects fees. eModal does appointment management but publishes no terminal counts, volumes or AI claims. If a proposal conflates these, the proposal has not been checked.
The Container Port Performance Index Does Not Settle This Question
At some point in an automation pitch, someone reaches for the World Bank and S&P Global Container Port Performance Index. It is the most authoritative-sounding artifact in the sector: 403 ports, more than 175,000 port calls, 247 million container moves, published by a multilateral institution alongside a major data provider.
It does not answer the automation question, and it is not close.
We extracted the full text of the report. The word stem "automat" appears eight times in the entire document, and there is no automated-versus-conventional comparison anywhere in it. The index measures vessel time in port, not moves per crane-hour, and it makes no attempt to attribute performance to terminal configuration. Its one substantive statement on the subject reads: "Even though automation alone does not magically double productivity, it significantly reduces variability and human error, making overall vessel handling times more predictable." That is a reasonable and modest claim — reduced variability, fewer human errors, more predictable handling — and it is a very different claim from a productivity uplift that amortises a capital programme.
For context on what the index does say: the top five ports on the 2024 data were Yangshan at 146.3, Fuzhou at 139.2, Port Said at 137.4, Dalian at 136.5 and Tanger-Med at 135.8. Useful for benchmarking vessel service. Useless for deciding whether to buy automated stacking cranes.
There is a broader lesson in this. The most-cited source in a sector is often the one least read. The CPPI gets quoted in automation discussions precisely because everyone recognises the name and almost nobody has opened the PDF. The same is true of the McKinsey study on the other side of the argument, which is why we stated its limitations in the same section as its findings rather than in a footnote.
Container OCR: An Accuracy Number Nobody Has Independently Tested
Optical character recognition at the gate and on the crane is the most mature machine-vision application in a terminal, and it is genuinely useful. It is also the clearest case in this industry of a number everybody quotes and nobody has verified.
Visy advertises "market leading accuracy of 97–99.5%" on its truck OCR portal — with no methodology, no per-attribute breakdown, and no named site. Camco's frequently quoted "98% guaranteed" sits on pages that block verification. Neither is dishonest on its face. Neither is checkable, and no independent third-party test of container-code OCR accuracy exists anywhere that we could find.
The most useful document in this whole area turns out to be a vendor-authored one that undercuts vendor marketing. Anton Bernaerd of Camco Technologies wrote a piece for Port Technology Internationaltitled, with commendable directness, "High OCR hit rates don't guarantee less exception rates." It is the closest thing to a methodology standard this field has, and it comes from someone selling the product.
- A hit rate needs 250 to 500 passages, with no images discarded: Bernaerd's stated threshold for a meaningful measurement. A demonstration over twenty containers, or over a set where unreadable images were quietly excluded, is not a hit rate — it is a highlight reel.
- The TOS auto-correction must be stripped out before scoring: Terminal operating systems correct OCR output against expected container lists. If you score the OCR after that correction, you are measuring the TOS, not the vision system. This single point invalidates a lot of field data.
- "Confidence level" is not a hit rate: A model reporting 99% confidence is reporting how sure it is, not how often it was right. These get conflated constantly, sometimes innocently. Ask which one the number is.
- Ask for a per-attribute breakdown: Container number, ISO code, check digit, door direction, seal presence, IMO placard and damage are separate recognition tasks with very different difficulty. A single blended figure hides the ones that will generate your exceptions.
- The exception path matters more than the hit rate: This is Bernaerd's actual argument. A 98% hit rate with a badly designed 2% exception path produces more clerical work than a 95% hit rate with a good one. Ask what happens on a miss, and time it.
For a sense of how much harder the adjacent problems are, look at the one task where independent academic measurement does exist. Published research in Science Progress in 2025 applied a YOLO-NAS detector to container damage detection and reported mean average precision of 91.2 percent, precision of 92.4 percent, and recall of 84.1 percent. Recall is the number that matters operationally: roughly one damaged container in six went undetected. That performance is genuinely useful as a first-pass screen that flags candidates for human inspection. It is nowhere near good enough to be the sole basis of a damage claim against a carrier or a trucker, and any vendor implying otherwise is selling you a dispute.
When a vendor quotes an accuracy figure, the useful follow-up is never "can you do better?" It is "how did you measure that?" A vendor with a real methodology will happily explain it. A vendor without one will offer you a bigger number.
— Frenchy Digital procurement principle
Where the Case Is Genuinely Strong: The Back Office, Not the Quay
Everything above is about equipment automation, because that is what the evidence base covers and what vendors mostly pitch. The work we actually recommend to terminal operators is somewhere else entirely, and it shares five properties: it is high-volume, document-heavy, rule-governed, currently done by people reading PDFs and emails, and completely independent of cranes, straddle carriers and labour agreements.
Nobody is going to write a press release about gate paperwork exception handling. That is exactly why it is available.
| Workflow | How it works today | What an agent does | How you measure it | Where the human stays |
|---|---|---|---|---|
| Gate paperwork exception handling | A clerk opens a mismatched EIR, booking or interchange document, compares it against the TOS record, and resolves or escalates | Extract fields from the document, compare against the TOS record, classify the exception type, draft the correction, route to a clerk for approval | Exceptions per shift, minutes per exception, share resolved without escalation | None. The container has already moved. |
| Booking and release mismatches | Email and phone chasing between carrier, forwarder and terminal to reconcile a release that will not clear | Read the thread, identify the specific field in conflict, assemble the evidence, draft the reply, hold the write until a human approves | Open mismatches, age of oldest, touches per resolution | Never auto-release. A release is a cargo-control decision. |
| Appointment administration | Staff manually rebooking, cancelling and reconciling appointment slots against vessel and yard reality | Draft rebooking proposals from the current slot picture, notify affected truckers, reconcile no-shows and duplicates | No-show rate, rebooking latency, slot utilisation | Slot allocation policy stays with operations. The agent proposes, it does not allocate. |
| Billing, demurrage and detention disputes | A dispute arrives as an email with attachments; someone reconstructs the timeline from three systems | Assemble the chronology from gate, yard and vessel records, cite the tariff clause, draft the response with the evidence attached | Days to first response, dispute cycle time, write-off rate | Credit and waiver decisions are human. Always. |
| Customs and regulatory documentation | Clerks re-key and cross-check declarations against manifests and bookings | Pre-fill, cross-check, flag inconsistencies before submission, and explain each flag | Rejection rate, rework hours, submission lead time | The declarant remains the legally responsible party. No autonomous filing. |
| Maintenance work-order triage | Fault reports and inspection notes arrive as free text; a planner sorts, dedupes and prioritises | Classify, deduplicate against open orders, propose a priority and a parts list, route to the planner | Time to triage, duplicate rate, backlog age | No safety-critical maintenance decision is automated. The planner decides. |
| Vessel and berth documentation packs | Assembling stowage confirmations, bay plans, damage notes and departure paperwork by hand | Assemble the pack, check completeness against a template, flag missing items | Pack completeness at first submission, assembly hours | Operational decisions on stowage and berthing stay with the planner. |
Terminal back-office workflows where AI agents have a defensible case — Frenchy Digital, 2026. Every row keeps a human on the decision that moves cargo or money.
Three things make this category different from the automation debate, and all three matter to how you fund it.
- The before-state is measurable in a fortnight: You can count gate exceptions per shift, minutes per exception, open booking mismatches and dispute cycle time from data you already have. You cannot easily construct a counterfactual for a $1.5bn crane programme. The entire evidence problem described above comes from the difficulty of measuring capital automation — and it evaporates when the unit of work is a document.
- The capital risk is two orders of magnitude smaller: A single-workflow agent is a $28k–$70k project on a four-to-nine-week clock. If it does not work, you have spent less than the feasibility study for a crane order and you have learned something specific about your own data.
- It does not touch the labour question: Gate clerks, billing analysts and maintenance planners are not covered by the automation clauses that make equipment projects politically expensive. Reducing the clerical load on a documentation team is a different conversation from reducing the manning on a crane, and pretending otherwise makes both conversations harder.
The TOS Write Path Is the Binding Constraint, Not Model Capability
Every project in this category succeeds or fails on the same thing, and it is not the model. It is whether the agent can write back into the terminal operating system.
Terminal operating systems are closed or semi-closed by design. Integration is typically EDI, flat-file exchange, or a restricted API whose scope depends on your modules, your version and your contract. Reading data out is usually solvable — reports, database views, message feeds. Writing a correction back is where projects stall: updating a booking, releasing a hold, amending a gate transaction, closing a work order. Sometimes the capability exists and is licensed separately. Sometimes it exists and the vendor will not support its use from outside their own product. Sometimes it does not exist at your version.
It is worth noting what the market leader actually publishes. Kaleris's Navis N4markets optimisation modules — yard inventory monitoring, tower checking, RTG and terminal-truck optimisation, a yard intelligence suite described as AI-augmented — and publishes no performance metrics for any of them. That is not a criticism of the product, which runs a very large share of the world's container terminals. It is a data point about the evidence environment: even the platform vendor is not putting numbers on the table.
| System | What integration actually looks like | What to do about it |
|---|---|---|
| Terminal operating system (Navis N4, Zodiac, Octopi and others) | Read is usually solvable — reports, database views, EDI feeds. Write is the hard problem and varies by module and contract. | Get the write path in writing from the TOS vendor before scoping. Assume nothing. |
| EDI messaging (CODECO, COPARN, COARRI, BAPLIE) | Mature, well-documented, and genuinely useful as an agent input. Also asynchronous, so state can be stale. | Treat EDI as an event source, not as truth. Reconcile against the TOS before acting. |
| Gate and OCR systems | Usually a vendor appliance with its own database and its own exception queue. Integration is a support ticket, not an API key. | Budget calendar time, not just engineering time |
| Appointment platforms | Third-party, per-terminal, with no common standard across San Pedro Bay or most other port complexes | Per-terminal integration effort, repeated. This is why unified systems keep being proposed and keep not existing. |
| Billing and ERP | Often the easiest integration in the stack, and the one with the clearest financial payback | Start here if you want a defensible first result |
| Email, EDM and shared drives | Where most of the actual exception work lives, and where untrusted external content enters | The highest-value input and the highest-risk one. Design the blast radius first. |
Integration surfaces in a container terminal and the practical constraint each imposes — Frenchy Digital, 2026.
Prompt injection, and why it is a terminal problem specifically
The highest-value workflows in the table above all involve reading untrusted external content: carrier emails, forwarder attachments, booking amendments, customs correspondence, dispute claims with PDFs attached. Any agent reading that content can be influenced by instructions embedded in it. This is not a theoretical risk and it is not solved — the honest framing is blast-radius reduction, not prevention.
The controls that work are architectural rather than linguistic. Treat retrieved content as data and never as instruction. Allowlist the tools the agent can call, per workflow. Reject any tool argument a human never supplied. Scope every lookup to the booking, container or work order already in context. Require explicit human approval for any write that moves cargo or money. Log every tool call with its arguments, so that a bad outcome is reconstructable. Map your design against the OWASP Top 10 for LLM Applications and the NIST AI Risk Management Framework.
If a vendor's answer to injection risk is that they tuned the system prompt, the answer is no. A system prompt is not a security boundary.
One further note on the operational context, because terminals have a specific reason to care about blast radius. The sector's recent cyber incidents were disruptive out of proportion to their sophistication: the Port of Nagoya was down roughly two days after a LockBit 3.0 attack in July 2023, and DP World Australia detected an incident on 10 November 2023, resumed operations by the 13th, and cleared a 30,137-container backlog by the 20th — with no ransomware found and no ransom demand. Terminals are tightly coupled systems where a two-day software outage becomes a two-week cargo problem. That is the right mental model for what an agent with excessive write access could do by accident.
Labour Is the Binding Constraint — and the Contract Language Is Not Public
No honest article about port automation can skip labour, and no honest article should overstate what it knows about the current agreements. Here is the line between the two.
What is confirmed about the most consequential recent dispute: the International Longshoremen's Association struck from 1 to 3 October 2024, involving roughly 47,000 workers across 36 East and Gulf Coast ports — the first such strike since 1977. The contract was extended to 15 January 2025 during negotiations. The union opened by demanding a 77 percent wage increase over six years and a complete ban on automation. The settlement was a six-year agreement with a 62 percent wage increase over its life, described as offering "full protections against automation" rather than the outright ban that had been demanded.
What the dispute does establish, regardless of the clause text, is the shape of the constraint. Equipment automation in a unionised port is a bargaining matter before it is an engineering matter, the negotiating positions are far apart, and the settlement pattern is expensive wage growth in exchange for constrained automation. Any capital business case that treats labour as a cost line rather than a counterparty has mispriced the project.
This is also the strongest structural argument for the back-office category. A gate-exception agent that helps a clerical team clear paperwork faster is not a manning question and does not enter that negotiation. That is not a loophole — it is a genuinely different kind of work, and conflating the two is what makes terminal AI conversations unnecessarily adversarial.
Capex Reality, and the Most Misread Number in Port Automation
The scale of the capital involved is why the evidence problem matters. These are not software budgets.
| Project or asset | Published cost | What it covers | Source quality |
|---|---|---|---|
| Long Beach Middle Harbor | US$1.49bn | Completed July 2021 — 69 electric stacking cranes, 102 battery AGVs, 18 dual-hoist quay cranes, 3.5m TEU | Port-reported |
| APM Terminals Maasvlakte II, Rotterdam | EUR 500m (2015), expansion >EUR 1bn (Mar 2023) | 130 ha, 2,000 m quay; up to 71 Lift AGVs added Oct 2024, fleet to exceed 140 | Operator-reported; only qualitative CO₂ and reliability claims published |
| Shanghai automated terminal | US$2.15bn | Recorded by the G20-backed Global Infrastructure Hub | Independently compiled project record |
| Port of Brisbane | A$250m | Global Infrastructure Hub project record | Independently compiled project record |
| Victoria International Container Terminal | A$650m | Global Infrastructure Hub project record | Independently compiled project record |
| Port of Virginia — Norfolk International Terminals North | > EUR 130m | 36 automated stacking cranes ordered May 2023, deliveries mid-2025 and mid-2027 | Vendor announcement; no productivity figures given |
| TraPac, Los Angeles | "Over $200 million" in equipment | Operator self-reported. The widely cited $510m figure is trade press and unverified — we do not use it. | Operator self-reported |
| Single ship-to-shore crane — the only verified price | R240m (about US$12.7m) | Liebherr STS at Transnet Port Elizabeth, April 2025 | The only single-crane price we could verify from a primary source |
| PEMA automation premium — READ THIS CAREFULLY | About US$1.1m more per crane | A CARMG versus an electrified RTG. This is the PREMIUM, not the crane price. | The most misread number in the sector |
Published capital figures for automated terminal programmes. Operator-reported figures are labelled as such; trade-press figures we could not verify have been excluded rather than reproduced.
On absolute crane pricing, the honest answer is that almost nothing verifiable is public. The only single-crane price we could confirm from a primary source is a Liebherr ship-to-shore crane supplied to Transnet at Port Elizabeth in April 2025 at R240 million, roughly US$12.7 million. The "average STS costs $11 million" and "ASCs start at $4 million" figures that circulate widely trace to market-research and vendor blogs with no underlying data, and we do not use them.
Predictive maintenance deserves a specific warning here, because it is the automation sub-pitch with the largest gap between enthusiasm and evidence. Konecranes made TRUCONNECT remote monitoring, CheckApp and its customer portal standard across all port cranesin April 2024. The announcement contains no quantified uptime, downtime or cost figure at all. Searches for quantified terminal predictive-maintenance case studies from the major equipment vendors returned nothing usable. If the original equipment manufacturer rolling this out fleet-wide will not publish a number, treat a reseller's number with corresponding suspicion.
Two other things worth knowing when a smart-port programme is invoked as proof of concept. The Port of Antwerp-Bruges APICA digital twin won a digitalisation award in 2025 and has published no measured results. Singapore's Maritime Digital Twin launched in March 2025 and was still being described as trialled by a US government trade brief in June 2026 — a pilot, not an operational system. The one substantive productivity claim in that programme comes from a different component: MPA's digitalPORT@SG, which consolidates sixteen clearance forms across more than 550 companies and is credited by MPA with saving up to 100,000 man-hours per year. Note what that is. It is a documentation and clearance system. It is back-office automation.
How to Test a Vendor Claim Before You Buy
Ten questions. Send them before the demo. They are derived directly from the failure patterns documented above, and each one is answerable by a vendor who has done the work.
| Question | Acceptable answer | Disqualifying answer |
|---|---|---|
| What is the baseline, and who measured it? | A named measurement window, a stated method, and someone other than the vendor or the terminal collecting the data | A percentage improvement with no before-state, or a before-state the vendor supplied |
| Is this an average or a peak? | A distribution, or at minimum a mean over a stated period | "Up to," which is a peak dressed as a norm |
| Which site produced this number? | A named terminal, a named year, a named configuration | "A major European terminal" — an unnamed site is an unfalsifiable claim |
| Has anyone outside the vendor reproduced it? | A published independent study, a customer willing to speak without the vendor on the call, or a trial you run yourself | A case study on the vendor's own site and nothing else |
| For OCR: how was the hit rate calculated? | 250 to 500 consecutive passages, no images discarded, terminal auto-correction stripped out, per-attribute breakdown | A single headline percentage. Ask specifically whether "confidence level" is being reported as a hit rate — it is not one. |
| What happens on an exception? | A measured exception rate and a documented human path | Only the success rate. A 98% hit rate with a badly designed 2% path is worse than 95% with a good one. |
| Can we run a two-week shadow trial on our data? | Yes, on historical exceptions, scored against what your staff actually decided | A demo environment with vendor-selected examples |
| What does the TOS write path look like at our version? | A specific answer naming modules, message types and any vendor dependency | "We integrate with all major TOS platforms" |
| Who is accountable when the agent is wrong? | A named human approval step before any write that moves cargo or money | "The model is 99% accurate," offered as a substitute for an approval step |
| Can we export the full audit trail? | Machine-readable export including tool calls, arguments, inputs and the human decision | A dashboard view you cannot extract or preserve |
Frenchy Digital due-diligence question set for terminal AI and automation procurement, 2026.
The shadow trial in row seven is worth more than the other nine combined. Take two weeks of your own historical exceptions — gate mismatches, booking conflicts, disputes — run them through the vendor's system, and score the output against what your staff actually decided at the time. It costs a fortnight and it converts every claim in the deck into a number from your own operation. A vendor who refuses is telling you something a reference call will not.
One structural point about evidence you generate yourself. The reason the independent literature on port automation is so thin is that nobody could construct a clean counterfactual for a capital programme — you cannot run the same terminal conventionally and automated in the same year. That constraint does not apply to a document workflow. You have a before-state, you have a holdout, and you can measure the difference in weeks. Use that advantage, because it is the one the equipment side never had.
Red Flags in a Port Automation or Terminal AI Pitch
Each of these corresponds to a specific documented failure in the evidence above. None of them are hypothetical.
| Red flag | Why it matters |
|---|---|
| A percentage with no baseline | Blue Visby's own trials illustrate the problem outside the terminal gate: the same vessel scored 28.2% against a 14-knot service speed and 7.9% against its 12-knot intended voyage speed. The number moved 20 points on baseline choice alone. |
| "Up to X%" anywhere in a business case | "Up to" is a peak observation. Peaks do not amortise capex. |
| The CPPI cited as proof automation works | The index makes no automated-versus-conventional comparison, and its own text warns that scores cannot be used to assess performance evolution over time. |
| A 99-point-something OCR accuracy figure | No independent test of container-code OCR exists. Ask how it was measured; the honest vendor answer is a methodology, not a bigger number. |
| A recycled statistic with no year attached | The "45% turn-time reduction" has been in circulation for nine years and originates in a terminal self-estimate. Ask when the number was produced and by whom. |
| Simulation or model output presented as measurement | Bureau Veritas validated the Blue Visby methodology for estimating effects — explicitly not measured savings. That distinction is routinely dropped in retelling. |
| Predictive maintenance ROI with no numbers | Konecranes made remote monitoring standard across all port cranes in April 2024 and the announcement contains no quantified uptime, downtime or cost figure. If the OEM will not publish one, be sceptical of a reseller who will. |
| An agent with autonomous write access to the TOS | A wrong autonomous release, hold or gate-out is a cargo-control incident, not a software bug. Human approval on every write that moves cargo or money. |
| No prompt-injection story for external email | Any agent reading carrier and forwarder correspondence is reading untrusted input. If the answer is "we tuned the prompt," the answer is no. |
| A crane price quoted from PEMA's automation figure | PEMA's roughly $1.1m per crane is the premium of a CARMG over an electrified RTG, not a crane price. This is the single most misread number in port automation. |
The Frenchy Digital red-flag list for terminal AI and automation buyers, 2026.
The baseline problem in the first row is worth one more sentence, because it generalises past the terminal gate. In documented voyage-optimisation trials, the same vessel scored a 28.2 percent saving against a 14-knot service speed and 7.9 percent against its 12-knot intended voyage speed. Same ship, same voyage, same software — a twenty-point swing produced entirely by which baseline the analyst chose. Whenever you are shown a percentage improvement without a stated baseline, the number is not wrong so much as it is meaningless.
A vendor who cannot tell you the baseline, the measurement window and who collected the data has not measured anything. They have observed something once and rounded it upward. That is not fraud, but it is not evidence either, and a capital committee should be able to tell the difference.
— Frenchy Digital buyer's principle
What It Costs to Build This Properly
These are the bands Frenchy Digital uses to scope terminal and port operations work in 2026. They assume the integration assessment and the measured before-state are in scope from day one, because those are what make the result defensible afterwards.
| Engagement | Range | Timeline | Typical scope |
|---|---|---|---|
| Discovery + workflow audit | $9k–$22k | 2–4 weeks | Workflow inventory, TOS and gate integration assessment, measured before-state for the top three candidate workflows, prioritised shortlist |
| Single-workflow agent (gate exceptions, appointment admin, work-order triage) | $28k–$70k | 4–9 weeks | One workflow end to end, human approval queue, audit logging, exception path, evaluation set built from your historical cases |
| Multi-workflow operations platform with system integration | $70k–$180k | 9–16 weeks | Several workflows, TOS and billing integration, EDI ingestion, shared document pipeline, role-based approvals, evals in CI |
| Enterprise / multi-terminal / regulated build | $180k–$420k+ | 14–24 weeks | Multi-site tenancy, full audit pipeline, human-in-the-loop controls, SOC 2 posture, disaster recovery testing, documentation package |
Frenchy Digital cost bands for port and terminal AI agent engagements, 2026.
Senior-led delivery runs $150 to $225 per hour, and ongoing retainers run $2,500 to $9,500 per month covering model and dependency upgrades, evaluation-set expansion, incident response and a quarterly technical review. Every engagement carries a 30-day post-launch warranty, and you receive a written scope with a fixed-price phased proposal within 5 business days of the discovery call.
A budgeting note that is specific to this sector. Compare these bands against the capital table above: the entire enterprise band is a rounding error against a single automated stacking crane order, let alone a $1.49bn terminal programme. That asymmetry is the actual argument for starting in the back office. You can run four workflow agents, measure all of them properly, and discard the two that did not earn their keep, for less than the cost of the consultancy work that would precede a capital decision.
Limitations and Honest Failure Modes
This article has spent nine sections holding vendors to an evidence standard. It would be indefensible not to apply the same standard to itself.
- The strongest finding here is old: McKinsey's productivity result comes from a 2017 opinion survey of 40-plus practitioners published in December 2018, and we found no replication. It is roughly seven and a half years old. Equipment, software and operating practice have all moved. We cite it because nothing published since contradicts it and because it had no commercial incentive to say what it said — not because it describes 2026.
- The turn-time data is one month from 2017: The Harbor Trucking Association GPS figures are genuine independent measurement, and they are a single month at a single port complex. Terminal mix, chassis availability, gate hours and volume all differ. TraPac at 114 minutes is a data point, not a verdict on automation.
- Two of the labour figures come from interested parties: The Economic Roundtable study is funded by the ILWU Coast Longshore Division and says so; the Prism Economics figures have a commissioner we could not verify. Moody's capacity estimate comes from a ratings agency and is disputed by the ITF. All three are labelled in the tables above rather than laundered into the narrative.
- We could not verify the operative ILA contract language: USMX's site errors and the ILA's links to a superseded agreement. Everything in the labour section is confined to what is confirmed about the strike and the settlement's headline terms. We make no claim about what the contract permits on semi-automation, and you should distrust anyone who does without citing the text.
- No independent OCR benchmark exists to check anyone against: That is a finding, but it is also a gap. It means we cannot tell you what accuracy is achievable in your gate configuration — only how to ask, and what a meaningful measurement would look like.
- Vendor deployment claims are almost entirely uncorroborated: Named installations exist; measured outcomes largely do not. Kaleris, Visy, Wabtec, Konecranes and the major smart-port programmes publish deployments, awards and module names, and almost no performance figures. We have reported the absence rather than filling it.
- Terminal predictive maintenance has essentially no published evidence: We looked. Konecranes' fleet-wide rollout announcement contains no quantified benefit, and no usable quantified case study from the major equipment vendors surfaced. If predictive maintenance is central to a proposal you are evaluating, treat the ROI model as entirely unvalidated.
- Back-office agents fail in ordinary software ways: Document extraction degrades on poor scans and unfamiliar formats. Exception classification drifts as carriers change their templates. Integration breaks on a TOS upgrade. None of these are exotic AI failures; all of them need monitoring, an owner and a maintenance budget, which is what the retainer is for.
- Prompt injection is unsolved: Any agent reading carrier or forwarder correspondence is processing untrusted input, and no current technique eliminates the risk. The design goal is blast-radius reduction — narrow tools, no unsupplied arguments, human approval on consequential writes, full tool-call logging — not immunity.
None of this argues against building. It argues for building the measurement alongside the agent, starting with one workflow where the before-state is countable, and being straight with your own capital committee about which numbers in the business case are measured and which are borrowed. The terminals that get value out of this are the ones that instrumented the before-state — and, notably, that is exactly the discipline the automation evidence base never managed.
The boundary holds regardless: these are administrative and documentation systems operating under human review. No agent described here releases cargo, allocates a berth, dispatches equipment or closes a safety-critical maintenance decision. A human approves every write that moves cargo or money, and the log proves it.
Evaluating an AI Pitch for Your Terminal?
Book a free 60-minute discovery call with Frenchy Digital — a senior-led Black-owned LA agency. You leave with an integration assessment, a measured before-state for one workflow, and a fixed-price phased proposal within 5 business days. Call +1 (424) 272-5601.
Evaluating an AI Pitch for Your Terminal?
Book a free 60-minute discovery call. You leave with an integration assessment, a measured before-state for one workflow, and a fixed-price phased proposal within 5 business days.
1517 S Bentley Ave Unit 204, Los Angeles CA 90025
Frequently Asked Questions
Sources & References
- 1McKinsey — The Future of Automated Ports (Dec 2018)↗
- 2OECD/ITF — Container Port Automation: Impacts and Implications (2021)↗
- 3World Bank / S&P Global — Container Port Performance Index 2024↗
- 4US EPA — GCT Bayonne's Drayage Truck Appointment System↗
- 5Transport Topics — Truck Turn Times Deteriorate at Ports of Los Angeles, Long Beach↗
- 6Port Technology International ed. 118 — Camco: High OCR Hit Rates Don't Guarantee Less Exception Rates (PDF)↗
- 7Science Progress (2025) — Container damage detection with YOLO-NAS (PMID 39887245)↗
- 8Kaleris — Navis N4 Terminal Operating System↗
- 9Global Infrastructure Hub — infrastructure project data↗
- 10PEMA — Port Equipment Manufacturers Association↗
- 11Visy Oy — container and truck OCR systems (vendor)↗
- 12Port of Long Beach — port information and Middle Harbor programme↗
- 13Port of Los Angeles — official port site↗
- 14Economic Roundtable — research publications (ILWU-funded automation study)↗
- 15IMO GreenVoyage2050 — Just-In-Time Arrivals study (MarineTraffic + EERA)↗
- 16OWASP Top 10 for LLM Applications↗
- 17NIST AI Risk Management Framework↗
- 18Konecranes — port crane monitoring and automation products↗

