Skip to main contentSkip to footer

    Top Rated & Verified

    Top Clutch App Development Company Black Owned United StatesTop Clutch Java Developers France 2026Top Clutch Service Line Blind Company Black Owned 2026Top Clutch App Development Company Minority Owned 2026Top Clutch Web Developers Black Owned 2026Top Clutch App Development Company Black Owned 2026Top Clutch Flutter Developers France 2026Top Clutch Health & Wellness App Developers France 2026Top Clutch Swift Company France 2026Top Clutch Machine Learning Company France 2026Top Clutch Chatbot Company France 2026Top Clutch Artificial Intelligence Company France 2026Top Clutch App Development Company Minority Owned Los Angeles
    Back to Blog
    Ports & Terminals
    August 9, 2026
    26 min read

    AI Agents for Port & TerminalOperations in 2026

    You are being pitched automation. This is what the neutral evidence actually found — McKinsey's practitioner survey, the OECD/ITF review, independent GPS turn-time data, and the statistics that turn out to be self-estimates — followed by the workflows where the case genuinely holds.

    AI agents in container port and marine terminal operations 2026 — gate exceptions, appointment administration, billing disputes and documentation workflows
    −7% to −15%
    Productivity change at automated ports, against an expected +10% to +35%
    McKinsey, The Future of Automated Ports (2018)
    ~4%
    Share of global container capacity that is automated
    OECD/ITF, Container Port Automation (2021)
    114 min
    Turn time at automated TraPac LA — the slowest terminal measured
    Harbor Trucking Association GPS data, July 2017
    $70k–$180k
    Multi-workflow terminal operations platform with system integration
    Frenchy Digital scoping, 2026

    Key Takeaways

    • McKinsey's practitioner survey found that against expectations of 25–55% opex reduction and 10–35% productivity gains, automated ports cut opex only 15–35% and productivity actually fell by 7–15%, with ROIC short of the ~8% industry norm by up to a percentage point. It is a 2017 opinion survey of 40+ practitioners, now about seven and a half years old, with no replication found — cite it for direction, not as a benchmark.
    • The OECD/ITF, an intergovernmental body with no product to sell, counted automated terminals at roughly 4% of global capacity and cited peer-reviewed work concluding that automation alone has no highly significant impact on terminal performance.
    • The cleanest vendor-versus-independent comparison available: ABB advertised a 45–55% labour-time reduction per crane; Oliveira and Varela measured 33%.
    • Independent GPS data contradicts the turn-time narrative. Harbor Trucking Association geofence data put the automated TraPac LA terminal at 114 minutes against 38 at Matson and 46 at Middle Harbor/LBCT — the automated terminal was the slowest measured.
    • The canonical "45% turn-time reduction" for truck appointment systems traces to a 2017 terminal self-estimate. The EPA's own page says "more than 40%" and attributes it explicitly to "GCT estimates," with no baseline and no methodology — and it has been recycled for nine years.
    • The World Bank / S&P Global Container Port Performance Index does not settle the automation question. The word stem "automat" appears eight times in the full report and there is no automated-versus-conventional comparison anywhere.
    • No independent third-party test of container-code OCR accuracy exists. Vendor figures of 97–99.5% carry no methodology, no per-attribute breakdown and no named site.
    • The genuinely strong case is back-office and documentation work — gate paperwork, booking and release exceptions, appointment administration, billing and demurrage disputes, customs documentation, maintenance work-order triage — none of which requires touching a crane or a union agreement.
    • Frenchy Digital cost bands: discovery $9k–$22k; single-workflow agent $28k–$70k; multi-workflow operations platform $70k–$180k; enterprise multi-terminal build $180k–$420k+.

    The Independent Evidence Contradicts the Sales Pitch

    Container terminals are unusual among the industries we write about, and the reason is worth stating up front. Most sectors being sold AI have almost no independent evidence about whether the technology delivers — you get vendor case studies, analyst reports commissioned by vendors, and a lot of adjectives. Ports are different. Terminal automation has been running at scale for over a decade, it has attracted attention from a global consultancy, an intergovernmental transport body, peer-reviewed academics, a trucking association with GPS data and the World Bank, and several of those parties have published.

    That evidence does not say what the sales deck says. That is the article.

    This is written for a terminal operator, a port authority executive or an operations director sitting across the table from a vendor. It is not an argument against automation, and it is certainly not an argument against AI in terminals. It is an argument that the specific claims being made to you have a documented track record of not surviving independent measurement, that you should know which claims those are, and that there is a separate category of work — unglamorous, back-office, document-heavy — where the case is genuinely strong and almost nobody is pitching it to you.

    The distinction that organises everything below. Equipment automation — automated stacking cranes, AGVs, remote-operated quay cranes — is a capital project measured in hundreds of millions, entangled with labour agreements, and supported by evidence that ranges from mixed to negative. Back-office agents — gate paperwork, booking exceptions, billing disputes, documentation — are a software project measured in tens of thousands, touch no union clause and no crane, and have measurable before-states. These are not the same purchase, and vendors benefit enormously from the confusion between them.

    One more framing note. Everything Frenchy Digital builds in this space is administrative and documentation automation under human review. Nothing below proposes an agent that moves a container, releases cargo, allocates a berth or makes a safety-critical maintenance call on its own. The evidence problem in this industry is bad enough without adding autonomy to it.

    McKinsey's Practitioner Survey: Expectations Versus Outcomes

    Start with the finding that a consultancy with a large logistics practice published about its own clients' industry, because it is the least likely thing to be marketing.

    McKinsey's "The Future of Automated Ports"surveyed industry practitioners about what they expected from automation and what they actually got. The gap is the story. Respondents expected automation to cut operating expenses by 25 to 55 percent and to raise productivity by 10 to 35 percent. In the report's own words, "today these expectations generally aren't realized, especially in fully automated projects." Operating expenses did fall — but only by 15 to 35 percent. And then the sentence the industry has never comfortably answered: "Worse, productivity actually falls, by 7 to 15 percent."

    MetricWhat practitioners expectedWhat the survey foundWhat it means for your business case
    Operating expense reduction25% to 55% lower15% to 35% lowerReal, but landing at the bottom of what the business case assumed
    Productivity10% to 35% higher7% to 15% LOWERThe sign flipped. This is the finding the sector has never satisfactorily answered.
    Return on invested capitalAt or above the industry norm of about 8%Short of that norm by up to one percentage pointA project that clears its hurdle rate on paper and misses it in operation
    What would justify the capexOpex 25% lower than a conventional terminal, or productivity up 30% while opex fell 10%Neither threshold was being met by the surveyed projects
    Moves per crane-hour"Low 20s" automated against "high 30s" conventionalExplicitly a single executive's anecdote in the source, not measured data. Do not use it as a benchmark.

    Expected versus realised outcomes at automated ports, per McKinsey's practitioner survey. Note the limitations stated below before citing any of this.

    The capital math in the same report is more damning than the headline. To justify the capital expenditure, McKinsey calculated, "operating expenses of an automated greenfield terminal would have to be 25 percent lower than those of a conventional one or productivity would have to rise by 30 percent while operating expenses fell by 10 percent." Neither condition was being met. Return on invested capital was "falling short by up to one percentage point from the industry norm of about 8 percent." A percentage point does not sound like much until you apply it to a $1.5bn programme over its life.

    Now the limitation, stated in the same breath, because anything else would be dishonest. This is an opinion survey of more than 40 industry practitioners, conducted in 2017 and published in December 2018. It reports what operators believed about their own projects, not audited terminal data. It is now about seven and a half years old, and we could not find a replication of it — nobody has run the survey again, and no comparable dataset has been published since. Equipment, software and operating practice have all moved in that time. Cite it for the direction and size of the expectation gap, which no later source contradicts. Do not cite it as a current performance benchmark.

    The same caution applies to the number people most want to quote from it. McKinsey mentions automated terminals running in the "low 20s" of moves per crane-hour against "high 30s" at conventional terminals. That comparison is presented in the source as one executive's anecdote. It is not measured data, it is not a survey result, and reproducing it as a benchmark — which happens constantly — is exactly the laundering problem this article exists to describe.

    A seven-year-old opinion survey with no replication is weak evidence about today. It is still stronger evidence than a vendor case study published last month, because it had no commercial reason to say what it said. Rank your sources by incentive before you rank them by recency.

    Frenchy Digital evaluation principle

    The OECD/ITF Review — and What Peer Review Actually Found

    The second source carries different weight, and it is the one we would put in front of a board. The OECD's International Transport Forum published "Container Port Automation: Impacts and Implications" in October 2021. The ITF is intergovernmental. It sells nothing, it rates nothing, and it has no client relationship with a terminal operator or an equipment manufacturer.

    Its first finding reframes the urgency. At the time of the review there were roughly 53 automated container terminals worldwide, representing about 4 percent of global capacity. After more than a decade of automation being described as the future of the sector, ninety-six percent of the world's container capacity is handled some other way. That is not an argument against automating. It is an argument against the specific pressure — everyone is doing this, you are falling behind — that shows up in the first five slides of every deck.

    Its second finding is the peer-reviewed one. The ITF cites Ghiara and Tei (2021), whose conclusion is that automation alone "cannot be considered to have a highly significant impact on port terminal performance." That is a null result from academic literature, not a consultancy opinion, and it is compatible with the McKinsey survey rather than contradicting it.

    FindingWhat the OECD/ITF review reportsHow to read it
    Scale of automationAbout 53 automated container terminals worldwide, roughly 4% of global capacityAutomation is a niche, not an industry standard. The pressure to keep up is smaller than the deck implies.
    Peer-reviewed assessmentGhiara and Tei (2021): automation alone "cannot be considered to have a highly significant impact on port terminal performance"The strongest available statement from peer review, and it is a null result
    APMT-Maasvlakte 2 remote quay cranes, 2019–2125 moves per crane-hour, with "hardly any week" exceeding 30 — lagging other Rotterdam terminalsA flagship automated terminal, measured over two years, underperforming its own port
    Qingdao43 moves per crane-hour reported in January 2020Automated terminals span an enormous range. The technology is not the variable that explains it.
    Labour reduction (Prism Economics, 2019)40–50% at TraPac LA, 50% at Patrick's Sydney, up to 85% at Qingdao[FUNDING UNVERIFIED — likely union-commissioned] Directionally consistent with automation cutting headcount, which nobody disputes
    Capacity claim (Moody's, 2019)Semi-automation at Norfolk International Terminal South would add 47% capacity on the same footprint[RATINGS AGENCY, NOT NEUTRAL] The ITF disputes the basis of the comparison
    Union-commissioned studyEconomic Roundtable, "Someone Else's Ocean" (June 2022): about 572 FTE per year eliminated at two automated terminals, 2020–21[FUNDED BY THE ILWU COAST LONGSHORE DIVISION] Read it as advocacy research with a real methodology, not as neutral measurement

    Findings from the OECD/ITF review of container port automation, 2021. Funding and commissioner labels are stated where they are known and flagged where they are not.

    The Rotterdam comparison in that table deserves its own paragraph, because it is the most specific operational number in the independent literature. APM Terminals' Maasvlakte 2 facility runs remote-operated quay cranes. Across 2019 to 2021, the ITF reports those cranes averaged 25 moves per crane-hour, with hardly any week exceeding 30 — lagging other terminals in the same port. Over the same period, Qingdao was reported at 43. Both are automated terminals. The eighteen-move spread between them tells you that automation is not the variable explaining terminal productivity, which is precisely what the peer-reviewed null result says.

    On labour numbers, and why we label the funding

    Two of the figures in that table come from studies with an interested commissioner, and we have labelled both. The Prism Economics headcount reductions — 40 to 50 percent at TraPac Los Angeles, 50 percent at Patrick's Sydney, up to 85 percent at Qingdao — appear in the ITF review, but we could not verify who commissioned the underlying work, and the likeliest answer is a union. The Economic Roundtable's "Someone Else's Ocean," which found roughly 572 full-time equivalents per year eliminated at two automated terminals in 2020–21, states its funder plainly: the ILWU Coast Longshore Division.

    Labelling them is not dismissing them. Union-funded research on job losses is often methodologically careful, and the direction of these findings is not seriously contested — automation reduces terminal headcount, which is the entire point of it. But a number produced by a party with a stake in the answer belongs in your analysis with the stake attached.

    The same rule applies in the other direction. Moody's 2019 estimate that semi-automation at Norfolk International Terminal South would add 47 percent capacity on the same footprint comes from a ratings agency, not a neutral observer, and the ITF disputes the basis of the comparison. We are not printing it as a fact either.

    Vendor Claim Versus Independent Measurement, Side by Side

    The cleanest way to show the pattern is to put the two columns next to each other. Every row below pairs something a vendor has published with whatever independent measurement exists for the same thing — and in several rows, the honest answer in the right-hand column is that no independent measurement exists at all.

    Claim areaWhat the vendor publishedWhat independent measurement showsThe gap
    Crane labour timeABB advertised a 45–55% reduction per crane [VENDOR]Oliveira and Varela (2017) measured 33%, as cited by the OECD/ITFThe advertised floor was 12 points above the measured result
    Container-code OCR accuracyVisy: "market leading accuracy of 97–99.5%" [VENDOR]; Camco: "98% guaranteed" [VENDOR, page not verifiable]No independent third-party test existsThere is nothing to compare the claim against. That is the finding.
    Quay crane throughputCamco/ZPMC at Beibu Gulf Port Qinzhou: "up to 36 moves per hour per crane," 8-second OCR per move [VENDOR PR via trade press]OECD/ITF on APMT-Maasvlakte 2: 25 moves per crane-hour sustained over 2019–21"Up to" is a peak. The independent figure is an average. They are not comparable.
    Truck turn time after an appointment system"45% reduction" at GCT Bayonne, widely recycled since 2017US EPA states "more than 40%" and attributes it to "GCT estimates" — no baseline, no methodologyThe canonical statistic in this category is a self-estimate
    Terminal capacity from TOS optimisationKaleris/Navis N4 named user Port of Helsingborg: "up to 30%" capacity gain in the existing area [VENDOR]No published metrics; Navis N4's own optimisation modules are marketed without AI claims or figuresA named customer is not a measurement
    Crane maintenance savings at Rotterdam"18% lower crane repair costs" and "20% shorter waits" circulate widelyTraceable only to SEO content farms; the Port of Rotterdam's own site frames its 4D digital twin as an aimREJECTED — we will not print these numbers as fact, and neither should your board deck

    Vendor claims paired with independent measurement where it exists. Vendor-sourced figures are labelled; unverifiable figures are marked as such rather than reproduced as fact.

    The first row is the cleanest example in the entire batch, and it is worth stating on its own. ABB advertised a 45 to 55 percent reduction in labour time per crane. Oliveira and Varela measured 33 percent. The measured result did not just miss the advertised range — it landed twelve points below the bottom of it. Notice also that 33 percent is a good result. A third of the labour time on a crane is a real saving that would justify a real investment. The problem is not that the technology did nothing. The problem is that the business case was built on a number nobody had measured.

    The last row is where we draw a line. The figures circulating about Rotterdam — 18 percent lower crane repair costs, 20 percent shorter waits — trace back only to content-farm articles with no primary source, while the Port of Rotterdam's own material describes its 4D digital twin as an aim. We will not print those numbers as facts. If they appear in a deck presented to you, ask for the primary source, and watch what happens.

    The pattern to internalise:vendor figures in this industry are not usually fabricated. They are usually peaks presented as averages, single sites presented as typical, simulations presented as measurements, or self-estimates presented as studies. The correction is not scepticism about the vendor's honesty. It is four questions — what is the baseline, is this a peak or a mean, which site produced it, and who measured it — asked every single time.

    Truck Turn Times: Independent GPS Data Against a Nine-Year-Old Self-Estimate

    Truck turn time is the metric drayage carriers care about most, it is the one appointment systems are sold on, and it is the one where independent measurement is available — because the Harbor Trucking Association collects GPS geofence data from its members' trucks rather than relying on terminal reporting.

    That data, for July 2017, put the San Pedro Bay average at 90 minutes, with 24 percent of trips exceeding two hours. The terminal-level breakdown is where it gets interesting.

    Terminal / scopeAverage turn timeNote
    TraPac, Los Angeles — automated114 minutesSlowest terminal in the dataset
    Middle Harbor / LBCT, Long Beach — automated46 minutesWell inside the average
    Matson38 minutesFastest terminal in the dataset
    Los Angeles, all terminals93 minutesPort-level average
    Long Beach, all terminals87 minutesPort-level average
    San Pedro Bay combined90 minutes24% of trips exceeded 120 minutes

    Harbor Trucking Association GPS geofence turn-time data, July 2017 — collected independently of terminal reporting.

    The automated TraPac Los Angeles terminal measured 114 minutes — the slowest in the dataset — against 38 minutes at Matson and 46 at Middle Harbor and LBCT. Be careful about how much weight this carries: it is one month of one association's data from 2017, terminals differ in cargo mix, chassis availability, gate hours and volume, and a single automated terminal at the bottom of one table is not proof that automation slows trucks down. But it is measurement, collected by a party with no interest in the automation debate, and it points the opposite way from the narrative.

    Now the appointment-system statistic, which is the more instructive story.

    The canonical number in this category is a self-estimate. The widely-cited "45 percent turn-time reduction" from a truck appointment system at New York/New Jersey traces to a 2017 terminal estimate reported in trade press. Go to the primary source and it gets thinner: the US EPA's own page on the GCT Bayonne system says turn times fell "more than 40%" and attributes that explicitly to "GCT estimates" — no baseline, no measurement window, no methodology. It has been recycled for nine years. Name that pattern when you see it, because this industry is full of it.

    None of this means appointment systems do not work. It means the number your vendor is quoting is not evidence that they work. And the state of play in the largest US port complex suggests the sector itself is not certain either: Los Angeles and Long Beach still have no unified appointment system across their twelve container terminals. In February 2023 the Harbor Trucking Association and the California Trucking Association formally asked for one interoperable system; the response was supportive and carried no timeline. As of the port's State of the Port address in January 2026, an $8 million California GO-Biz grant is still being used to extend the Port of Los Angeles's Universal Truck Appointment System to Long Beach terminals — in progress, with no published measured turn-time result.

    Two clarifications that will save you an embarrassing slide. PortCheck is not an appointment system: per its own description it was founded in 2008 as an online rate-collection service for the Los Angeles/Long Beach Clean Truck Fund, and it collects fees. eModal does appointment management but publishes no terminal counts, volumes or AI claims. If a proposal conflates these, the proposal has not been checked.

    The Container Port Performance Index Does Not Settle This Question

    At some point in an automation pitch, someone reaches for the World Bank and S&P Global Container Port Performance Index. It is the most authoritative-sounding artifact in the sector: 403 ports, more than 175,000 port calls, 247 million container moves, published by a multilateral institution alongside a major data provider.

    It does not answer the automation question, and it is not close.

    We extracted the full text of the report. The word stem "automat" appears eight times in the entire document, and there is no automated-versus-conventional comparison anywhere in it. The index measures vessel time in port, not moves per crane-hour, and it makes no attempt to attribute performance to terminal configuration. Its one substantive statement on the subject reads: "Even though automation alone does not magically double productivity, it significantly reduces variability and human error, making overall vessel handling times more predictable." That is a reasonable and modest claim — reduced variability, fewer human errors, more predictable handling — and it is a very different claim from a productivity uplift that amortises a capital programme.

    Say this plainly to anyone who cites the CPPI at you:the index is a ranking of ports by vessel time in port. It is not a study of automation, it contains no automated-versus-conventional analysis, and its own text carries the caveat that "the scores cannot be used to assess performance evolution over time" because they are re-normalised annually. Anyone presenting it as evidence for or against automation ROI is overreaching — whichever side of the argument they are on.

    For context on what the index does say: the top five ports on the 2024 data were Yangshan at 146.3, Fuzhou at 139.2, Port Said at 137.4, Dalian at 136.5 and Tanger-Med at 135.8. Useful for benchmarking vessel service. Useless for deciding whether to buy automated stacking cranes.

    There is a broader lesson in this. The most-cited source in a sector is often the one least read. The CPPI gets quoted in automation discussions precisely because everyone recognises the name and almost nobody has opened the PDF. The same is true of the McKinsey study on the other side of the argument, which is why we stated its limitations in the same section as its findings rather than in a footnote.

    Container OCR: An Accuracy Number Nobody Has Independently Tested

    Optical character recognition at the gate and on the crane is the most mature machine-vision application in a terminal, and it is genuinely useful. It is also the clearest case in this industry of a number everybody quotes and nobody has verified.

    Visy advertises "market leading accuracy of 97–99.5%" on its truck OCR portal — with no methodology, no per-attribute breakdown, and no named site. Camco's frequently quoted "98% guaranteed" sits on pages that block verification. Neither is dishonest on its face. Neither is checkable, and no independent third-party test of container-code OCR accuracy exists anywhere that we could find.

    The most useful document in this whole area turns out to be a vendor-authored one that undercuts vendor marketing. Anton Bernaerd of Camco Technologies wrote a piece for Port Technology Internationaltitled, with commendable directness, "High OCR hit rates don't guarantee less exception rates." It is the closest thing to a methodology standard this field has, and it comes from someone selling the product.

    • A hit rate needs 250 to 500 passages, with no images discarded: Bernaerd's stated threshold for a meaningful measurement. A demonstration over twenty containers, or over a set where unreadable images were quietly excluded, is not a hit rate — it is a highlight reel.
    • The TOS auto-correction must be stripped out before scoring: Terminal operating systems correct OCR output against expected container lists. If you score the OCR after that correction, you are measuring the TOS, not the vision system. This single point invalidates a lot of field data.
    • "Confidence level" is not a hit rate: A model reporting 99% confidence is reporting how sure it is, not how often it was right. These get conflated constantly, sometimes innocently. Ask which one the number is.
    • Ask for a per-attribute breakdown: Container number, ISO code, check digit, door direction, seal presence, IMO placard and damage are separate recognition tasks with very different difficulty. A single blended figure hides the ones that will generate your exceptions.
    • The exception path matters more than the hit rate: This is Bernaerd's actual argument. A 98% hit rate with a badly designed 2% exception path produces more clerical work than a 95% hit rate with a good one. Ask what happens on a miss, and time it.

    For a sense of how much harder the adjacent problems are, look at the one task where independent academic measurement does exist. Published research in Science Progress in 2025 applied a YOLO-NAS detector to container damage detection and reported mean average precision of 91.2 percent, precision of 92.4 percent, and recall of 84.1 percent. Recall is the number that matters operationally: roughly one damaged container in six went undetected. That performance is genuinely useful as a first-pass screen that flags candidates for human inspection. It is nowhere near good enough to be the sole basis of a damage claim against a carrier or a trucker, and any vendor implying otherwise is selling you a dispute.

    When a vendor quotes an accuracy figure, the useful follow-up is never "can you do better?" It is "how did you measure that?" A vendor with a real methodology will happily explain it. A vendor without one will offer you a bigger number.

    Frenchy Digital procurement principle

    Where the Case Is Genuinely Strong: The Back Office, Not the Quay

    Everything above is about equipment automation, because that is what the evidence base covers and what vendors mostly pitch. The work we actually recommend to terminal operators is somewhere else entirely, and it shares five properties: it is high-volume, document-heavy, rule-governed, currently done by people reading PDFs and emails, and completely independent of cranes, straddle carriers and labour agreements.

    Nobody is going to write a press release about gate paperwork exception handling. That is exactly why it is available.

    WorkflowHow it works todayWhat an agent doesHow you measure itWhere the human stays
    Gate paperwork exception handlingA clerk opens a mismatched EIR, booking or interchange document, compares it against the TOS record, and resolves or escalatesExtract fields from the document, compare against the TOS record, classify the exception type, draft the correction, route to a clerk for approvalExceptions per shift, minutes per exception, share resolved without escalationNone. The container has already moved.
    Booking and release mismatchesEmail and phone chasing between carrier, forwarder and terminal to reconcile a release that will not clearRead the thread, identify the specific field in conflict, assemble the evidence, draft the reply, hold the write until a human approvesOpen mismatches, age of oldest, touches per resolutionNever auto-release. A release is a cargo-control decision.
    Appointment administrationStaff manually rebooking, cancelling and reconciling appointment slots against vessel and yard realityDraft rebooking proposals from the current slot picture, notify affected truckers, reconcile no-shows and duplicatesNo-show rate, rebooking latency, slot utilisationSlot allocation policy stays with operations. The agent proposes, it does not allocate.
    Billing, demurrage and detention disputesA dispute arrives as an email with attachments; someone reconstructs the timeline from three systemsAssemble the chronology from gate, yard and vessel records, cite the tariff clause, draft the response with the evidence attachedDays to first response, dispute cycle time, write-off rateCredit and waiver decisions are human. Always.
    Customs and regulatory documentationClerks re-key and cross-check declarations against manifests and bookingsPre-fill, cross-check, flag inconsistencies before submission, and explain each flagRejection rate, rework hours, submission lead timeThe declarant remains the legally responsible party. No autonomous filing.
    Maintenance work-order triageFault reports and inspection notes arrive as free text; a planner sorts, dedupes and prioritisesClassify, deduplicate against open orders, propose a priority and a parts list, route to the plannerTime to triage, duplicate rate, backlog ageNo safety-critical maintenance decision is automated. The planner decides.
    Vessel and berth documentation packsAssembling stowage confirmations, bay plans, damage notes and departure paperwork by handAssemble the pack, check completeness against a template, flag missing itemsPack completeness at first submission, assembly hoursOperational decisions on stowage and berthing stay with the planner.

    Terminal back-office workflows where AI agents have a defensible case — Frenchy Digital, 2026. Every row keeps a human on the decision that moves cargo or money.

    Three things make this category different from the automation debate, and all three matter to how you fund it.

    • The before-state is measurable in a fortnight: You can count gate exceptions per shift, minutes per exception, open booking mismatches and dispute cycle time from data you already have. You cannot easily construct a counterfactual for a $1.5bn crane programme. The entire evidence problem described above comes from the difficulty of measuring capital automation — and it evaporates when the unit of work is a document.
    • The capital risk is two orders of magnitude smaller: A single-workflow agent is a $28k–$70k project on a four-to-nine-week clock. If it does not work, you have spent less than the feasibility study for a crane order and you have learned something specific about your own data.
    • It does not touch the labour question: Gate clerks, billing analysts and maintenance planners are not covered by the automation clauses that make equipment projects politically expensive. Reducing the clerical load on a documentation team is a different conversation from reducing the manning on a crane, and pretending otherwise makes both conversations harder.
    The honest version of the pitch: we do not know whether automating your quay cranes will raise productivity, and neither does anyone else, because the independent evidence is mixed to negative and seven years stale. We do know that if your billing team spends nine hours a week reconstructing demurrage timelines from three systems, an agent that assembles the chronology and drafts the response with evidence attached will give most of those hours back — and you can measure it in the first month.

    The TOS Write Path Is the Binding Constraint, Not Model Capability

    Every project in this category succeeds or fails on the same thing, and it is not the model. It is whether the agent can write back into the terminal operating system.

    Terminal operating systems are closed or semi-closed by design. Integration is typically EDI, flat-file exchange, or a restricted API whose scope depends on your modules, your version and your contract. Reading data out is usually solvable — reports, database views, message feeds. Writing a correction back is where projects stall: updating a booking, releasing a hold, amending a gate transaction, closing a work order. Sometimes the capability exists and is licensed separately. Sometimes it exists and the vendor will not support its use from outside their own product. Sometimes it does not exist at your version.

    It is worth noting what the market leader actually publishes. Kaleris's Navis N4markets optimisation modules — yard inventory monitoring, tower checking, RTG and terminal-truck optimisation, a yard intelligence suite described as AI-augmented — and publishes no performance metrics for any of them. That is not a criticism of the product, which runs a very large share of the world's container terminals. It is a data point about the evidence environment: even the platform vendor is not putting numbers on the table.

    SystemWhat integration actually looks likeWhat to do about it
    Terminal operating system (Navis N4, Zodiac, Octopi and others)Read is usually solvable — reports, database views, EDI feeds. Write is the hard problem and varies by module and contract.Get the write path in writing from the TOS vendor before scoping. Assume nothing.
    EDI messaging (CODECO, COPARN, COARRI, BAPLIE)Mature, well-documented, and genuinely useful as an agent input. Also asynchronous, so state can be stale.Treat EDI as an event source, not as truth. Reconcile against the TOS before acting.
    Gate and OCR systemsUsually a vendor appliance with its own database and its own exception queue. Integration is a support ticket, not an API key.Budget calendar time, not just engineering time
    Appointment platformsThird-party, per-terminal, with no common standard across San Pedro Bay or most other port complexesPer-terminal integration effort, repeated. This is why unified systems keep being proposed and keep not existing.
    Billing and ERPOften the easiest integration in the stack, and the one with the clearest financial paybackStart here if you want a defensible first result
    Email, EDM and shared drivesWhere most of the actual exception work lives, and where untrusted external content entersThe highest-value input and the highest-risk one. Design the blast radius first.

    Integration surfaces in a container terminal and the practical constraint each imposes — Frenchy Digital, 2026.

    Prompt injection, and why it is a terminal problem specifically

    The highest-value workflows in the table above all involve reading untrusted external content: carrier emails, forwarder attachments, booking amendments, customs correspondence, dispute claims with PDFs attached. Any agent reading that content can be influenced by instructions embedded in it. This is not a theoretical risk and it is not solved — the honest framing is blast-radius reduction, not prevention.

    The controls that work are architectural rather than linguistic. Treat retrieved content as data and never as instruction. Allowlist the tools the agent can call, per workflow. Reject any tool argument a human never supplied. Scope every lookup to the booking, container or work order already in context. Require explicit human approval for any write that moves cargo or money. Log every tool call with its arguments, so that a bad outcome is reconstructable. Map your design against the OWASP Top 10 for LLM Applications and the NIST AI Risk Management Framework.

    If a vendor's answer to injection risk is that they tuned the system prompt, the answer is no. A system prompt is not a security boundary.

    One further note on the operational context, because terminals have a specific reason to care about blast radius. The sector's recent cyber incidents were disruptive out of proportion to their sophistication: the Port of Nagoya was down roughly two days after a LockBit 3.0 attack in July 2023, and DP World Australia detected an incident on 10 November 2023, resumed operations by the 13th, and cleared a 30,137-container backlog by the 20th — with no ransomware found and no ransom demand. Terminals are tightly coupled systems where a two-day software outage becomes a two-week cargo problem. That is the right mental model for what an agent with excessive write access could do by accident.

    Labour Is the Binding Constraint — and the Contract Language Is Not Public

    No honest article about port automation can skip labour, and no honest article should overstate what it knows about the current agreements. Here is the line between the two.

    What is confirmed about the most consequential recent dispute: the International Longshoremen's Association struck from 1 to 3 October 2024, involving roughly 47,000 workers across 36 East and Gulf Coast ports — the first such strike since 1977. The contract was extended to 15 January 2025 during negotiations. The union opened by demanding a 77 percent wage increase over six years and a complete ban on automation. The settlement was a six-year agreement with a 62 percent wage increase over its life, described as offering "full protections against automation" rather than the outright ban that had been demanded.

    What we will not tell you, and why.The operative contract language on semi-automation is not publicly verifiable. The USMX site returns an error and the ILA site links only to the superseded 2018–2024 agreement. That means we cannot responsibly characterise what the current contract permits or forbids regarding rail-mounted gantry cranes, manning conditions, or what counts as "semi-automated" — and neither can any vendor, consultant or article that tells you otherwise without citing the text. Ask your labour counsel and read your own copy. Treat a confident claim about that language, from anyone, as a signal about their standards.

    What the dispute does establish, regardless of the clause text, is the shape of the constraint. Equipment automation in a unionised port is a bargaining matter before it is an engineering matter, the negotiating positions are far apart, and the settlement pattern is expensive wage growth in exchange for constrained automation. Any capital business case that treats labour as a cost line rather than a counterparty has mispriced the project.

    This is also the strongest structural argument for the back-office category. A gate-exception agent that helps a clerical team clear paperwork faster is not a manning question and does not enter that negotiation. That is not a loophole — it is a genuinely different kind of work, and conflating the two is what makes terminal AI conversations unnecessarily adversarial.

    Capex Reality, and the Most Misread Number in Port Automation

    The scale of the capital involved is why the evidence problem matters. These are not software budgets.

    Project or assetPublished costWhat it coversSource quality
    Long Beach Middle HarborUS$1.49bnCompleted July 2021 — 69 electric stacking cranes, 102 battery AGVs, 18 dual-hoist quay cranes, 3.5m TEUPort-reported
    APM Terminals Maasvlakte II, RotterdamEUR 500m (2015), expansion >EUR 1bn (Mar 2023)130 ha, 2,000 m quay; up to 71 Lift AGVs added Oct 2024, fleet to exceed 140Operator-reported; only qualitative CO₂ and reliability claims published
    Shanghai automated terminalUS$2.15bnRecorded by the G20-backed Global Infrastructure HubIndependently compiled project record
    Port of BrisbaneA$250mGlobal Infrastructure Hub project recordIndependently compiled project record
    Victoria International Container TerminalA$650mGlobal Infrastructure Hub project recordIndependently compiled project record
    Port of Virginia — Norfolk International Terminals North> EUR 130m36 automated stacking cranes ordered May 2023, deliveries mid-2025 and mid-2027Vendor announcement; no productivity figures given
    TraPac, Los Angeles"Over $200 million" in equipmentOperator self-reported. The widely cited $510m figure is trade press and unverified — we do not use it.Operator self-reported
    Single ship-to-shore crane — the only verified priceR240m (about US$12.7m)Liebherr STS at Transnet Port Elizabeth, April 2025The only single-crane price we could verify from a primary source
    PEMA automation premium — READ THIS CAREFULLYAbout US$1.1m more per craneA CARMG versus an electrified RTG. This is the PREMIUM, not the crane price.The most misread number in the sector

    Published capital figures for automated terminal programmes. Operator-reported figures are labelled as such; trade-press figures we could not verify have been excluded rather than reproduced.

    The single most misread number in this sector. PEMA, the Port Equipment Manufacturers Association, publishes a figure of roughly US$1.1 million per crane in the context of automation. That is not the price of a crane. It is the premium — the additional cost of an automated rail-mounted gantry over an electrified rubber-tyred gantry. We have seen it quoted as a crane price in serious documents. If you see a per-crane automation cost that looks implausibly low, this is almost always why.

    On absolute crane pricing, the honest answer is that almost nothing verifiable is public. The only single-crane price we could confirm from a primary source is a Liebherr ship-to-shore crane supplied to Transnet at Port Elizabeth in April 2025 at R240 million, roughly US$12.7 million. The "average STS costs $11 million" and "ASCs start at $4 million" figures that circulate widely trace to market-research and vendor blogs with no underlying data, and we do not use them.

    Predictive maintenance deserves a specific warning here, because it is the automation sub-pitch with the largest gap between enthusiasm and evidence. Konecranes made TRUCONNECT remote monitoring, CheckApp and its customer portal standard across all port cranesin April 2024. The announcement contains no quantified uptime, downtime or cost figure at all. Searches for quantified terminal predictive-maintenance case studies from the major equipment vendors returned nothing usable. If the original equipment manufacturer rolling this out fleet-wide will not publish a number, treat a reseller's number with corresponding suspicion.

    Two other things worth knowing when a smart-port programme is invoked as proof of concept. The Port of Antwerp-Bruges APICA digital twin won a digitalisation award in 2025 and has published no measured results. Singapore's Maritime Digital Twin launched in March 2025 and was still being described as trialled by a US government trade brief in June 2026 — a pilot, not an operational system. The one substantive productivity claim in that programme comes from a different component: MPA's digitalPORT@SG, which consolidates sixteen clearance forms across more than 550 companies and is credited by MPA with saving up to 100,000 man-hours per year. Note what that is. It is a documentation and clearance system. It is back-office automation.

    How to Test a Vendor Claim Before You Buy

    Ten questions. Send them before the demo. They are derived directly from the failure patterns documented above, and each one is answerable by a vendor who has done the work.

    QuestionAcceptable answerDisqualifying answer
    What is the baseline, and who measured it?A named measurement window, a stated method, and someone other than the vendor or the terminal collecting the dataA percentage improvement with no before-state, or a before-state the vendor supplied
    Is this an average or a peak?A distribution, or at minimum a mean over a stated period"Up to," which is a peak dressed as a norm
    Which site produced this number?A named terminal, a named year, a named configuration"A major European terminal" — an unnamed site is an unfalsifiable claim
    Has anyone outside the vendor reproduced it?A published independent study, a customer willing to speak without the vendor on the call, or a trial you run yourselfA case study on the vendor's own site and nothing else
    For OCR: how was the hit rate calculated?250 to 500 consecutive passages, no images discarded, terminal auto-correction stripped out, per-attribute breakdownA single headline percentage. Ask specifically whether "confidence level" is being reported as a hit rate — it is not one.
    What happens on an exception?A measured exception rate and a documented human pathOnly the success rate. A 98% hit rate with a badly designed 2% path is worse than 95% with a good one.
    Can we run a two-week shadow trial on our data?Yes, on historical exceptions, scored against what your staff actually decidedA demo environment with vendor-selected examples
    What does the TOS write path look like at our version?A specific answer naming modules, message types and any vendor dependency"We integrate with all major TOS platforms"
    Who is accountable when the agent is wrong?A named human approval step before any write that moves cargo or money"The model is 99% accurate," offered as a substitute for an approval step
    Can we export the full audit trail?Machine-readable export including tool calls, arguments, inputs and the human decisionA dashboard view you cannot extract or preserve

    Frenchy Digital due-diligence question set for terminal AI and automation procurement, 2026.

    The shadow trial in row seven is worth more than the other nine combined. Take two weeks of your own historical exceptions — gate mismatches, booking conflicts, disputes — run them through the vendor's system, and score the output against what your staff actually decided at the time. It costs a fortnight and it converts every claim in the deck into a number from your own operation. A vendor who refuses is telling you something a reference call will not.

    One structural point about evidence you generate yourself. The reason the independent literature on port automation is so thin is that nobody could construct a clean counterfactual for a capital programme — you cannot run the same terminal conventionally and automated in the same year. That constraint does not apply to a document workflow. You have a before-state, you have a holdout, and you can measure the difference in weeks. Use that advantage, because it is the one the equipment side never had.

    Red Flags in a Port Automation or Terminal AI Pitch

    Each of these corresponds to a specific documented failure in the evidence above. None of them are hypothetical.

    Red flagWhy it matters
    A percentage with no baselineBlue Visby's own trials illustrate the problem outside the terminal gate: the same vessel scored 28.2% against a 14-knot service speed and 7.9% against its 12-knot intended voyage speed. The number moved 20 points on baseline choice alone.
    "Up to X%" anywhere in a business case"Up to" is a peak observation. Peaks do not amortise capex.
    The CPPI cited as proof automation worksThe index makes no automated-versus-conventional comparison, and its own text warns that scores cannot be used to assess performance evolution over time.
    A 99-point-something OCR accuracy figureNo independent test of container-code OCR exists. Ask how it was measured; the honest vendor answer is a methodology, not a bigger number.
    A recycled statistic with no year attachedThe "45% turn-time reduction" has been in circulation for nine years and originates in a terminal self-estimate. Ask when the number was produced and by whom.
    Simulation or model output presented as measurementBureau Veritas validated the Blue Visby methodology for estimating effects — explicitly not measured savings. That distinction is routinely dropped in retelling.
    Predictive maintenance ROI with no numbersKonecranes made remote monitoring standard across all port cranes in April 2024 and the announcement contains no quantified uptime, downtime or cost figure. If the OEM will not publish one, be sceptical of a reseller who will.
    An agent with autonomous write access to the TOSA wrong autonomous release, hold or gate-out is a cargo-control incident, not a software bug. Human approval on every write that moves cargo or money.
    No prompt-injection story for external emailAny agent reading carrier and forwarder correspondence is reading untrusted input. If the answer is "we tuned the prompt," the answer is no.
    A crane price quoted from PEMA's automation figurePEMA's roughly $1.1m per crane is the premium of a CARMG over an electrified RTG, not a crane price. This is the single most misread number in port automation.

    The Frenchy Digital red-flag list for terminal AI and automation buyers, 2026.

    The baseline problem in the first row is worth one more sentence, because it generalises past the terminal gate. In documented voyage-optimisation trials, the same vessel scored a 28.2 percent saving against a 14-knot service speed and 7.9 percent against its 12-knot intended voyage speed. Same ship, same voyage, same software — a twenty-point swing produced entirely by which baseline the analyst chose. Whenever you are shown a percentage improvement without a stated baseline, the number is not wrong so much as it is meaningless.

    A vendor who cannot tell you the baseline, the measurement window and who collected the data has not measured anything. They have observed something once and rounded it upward. That is not fraud, but it is not evidence either, and a capital committee should be able to tell the difference.

    Frenchy Digital buyer's principle

    What It Costs to Build This Properly

    These are the bands Frenchy Digital uses to scope terminal and port operations work in 2026. They assume the integration assessment and the measured before-state are in scope from day one, because those are what make the result defensible afterwards.

    EngagementRangeTimelineTypical scope
    Discovery + workflow audit$9k–$22k2–4 weeksWorkflow inventory, TOS and gate integration assessment, measured before-state for the top three candidate workflows, prioritised shortlist
    Single-workflow agent (gate exceptions, appointment admin, work-order triage)$28k–$70k4–9 weeksOne workflow end to end, human approval queue, audit logging, exception path, evaluation set built from your historical cases
    Multi-workflow operations platform with system integration$70k–$180k9–16 weeksSeveral workflows, TOS and billing integration, EDI ingestion, shared document pipeline, role-based approvals, evals in CI
    Enterprise / multi-terminal / regulated build$180k–$420k+14–24 weeksMulti-site tenancy, full audit pipeline, human-in-the-loop controls, SOC 2 posture, disaster recovery testing, documentation package

    Frenchy Digital cost bands for port and terminal AI agent engagements, 2026.

    Senior-led delivery runs $150 to $225 per hour, and ongoing retainers run $2,500 to $9,500 per month covering model and dependency upgrades, evaluation-set expansion, incident response and a quarterly technical review. Every engagement carries a 30-day post-launch warranty, and you receive a written scope with a fixed-price phased proposal within 5 business days of the discovery call.

    Included at every tier: the integration assessment covering your TOS write path, a measured before-state for the target workflow, human approval on every write that moves cargo or money, audit logging with tool-call capture, an evaluation set built from your own historical exceptions, and full source-code and IP ownership transferred to you at delivery. Frenchy Digital is a senior-led Black-owned Los Angeles agency, which matters if your port authority or terminal group runs a supplier-diversity programme — and we do not build lock-in.

    A budgeting note that is specific to this sector. Compare these bands against the capital table above: the entire enterprise band is a rounding error against a single automated stacking crane order, let alone a $1.49bn terminal programme. That asymmetry is the actual argument for starting in the back office. You can run four workflow agents, measure all of them properly, and discard the two that did not earn their keep, for less than the cost of the consultancy work that would precede a capital decision.

    Limitations and Honest Failure Modes

    This article has spent nine sections holding vendors to an evidence standard. It would be indefensible not to apply the same standard to itself.

    • The strongest finding here is old: McKinsey's productivity result comes from a 2017 opinion survey of 40-plus practitioners published in December 2018, and we found no replication. It is roughly seven and a half years old. Equipment, software and operating practice have all moved. We cite it because nothing published since contradicts it and because it had no commercial incentive to say what it said — not because it describes 2026.
    • The turn-time data is one month from 2017: The Harbor Trucking Association GPS figures are genuine independent measurement, and they are a single month at a single port complex. Terminal mix, chassis availability, gate hours and volume all differ. TraPac at 114 minutes is a data point, not a verdict on automation.
    • Two of the labour figures come from interested parties: The Economic Roundtable study is funded by the ILWU Coast Longshore Division and says so; the Prism Economics figures have a commissioner we could not verify. Moody's capacity estimate comes from a ratings agency and is disputed by the ITF. All three are labelled in the tables above rather than laundered into the narrative.
    • We could not verify the operative ILA contract language: USMX's site errors and the ILA's links to a superseded agreement. Everything in the labour section is confined to what is confirmed about the strike and the settlement's headline terms. We make no claim about what the contract permits on semi-automation, and you should distrust anyone who does without citing the text.
    • No independent OCR benchmark exists to check anyone against: That is a finding, but it is also a gap. It means we cannot tell you what accuracy is achievable in your gate configuration — only how to ask, and what a meaningful measurement would look like.
    • Vendor deployment claims are almost entirely uncorroborated: Named installations exist; measured outcomes largely do not. Kaleris, Visy, Wabtec, Konecranes and the major smart-port programmes publish deployments, awards and module names, and almost no performance figures. We have reported the absence rather than filling it.
    • Terminal predictive maintenance has essentially no published evidence: We looked. Konecranes' fleet-wide rollout announcement contains no quantified benefit, and no usable quantified case study from the major equipment vendors surfaced. If predictive maintenance is central to a proposal you are evaluating, treat the ROI model as entirely unvalidated.
    • Back-office agents fail in ordinary software ways: Document extraction degrades on poor scans and unfamiliar formats. Exception classification drifts as carriers change their templates. Integration breaks on a TOS upgrade. None of these are exotic AI failures; all of them need monitoring, an owner and a maintenance budget, which is what the retainer is for.
    • Prompt injection is unsolved: Any agent reading carrier or forwarder correspondence is processing untrusted input, and no current technique eliminates the risk. The design goal is blast-radius reduction — narrow tools, no unsupplied arguments, human approval on consequential writes, full tool-call logging — not immunity.

    None of this argues against building. It argues for building the measurement alongside the agent, starting with one workflow where the before-state is countable, and being straight with your own capital committee about which numbers in the business case are measured and which are borrowed. The terminals that get value out of this are the ones that instrumented the before-state — and, notably, that is exactly the discipline the automation evidence base never managed.

    The boundary holds regardless: these are administrative and documentation systems operating under human review. No agent described here releases cargo, allocates a berth, dispatches equipment or closes a safety-critical maintenance decision. A human approves every write that moves cargo or money, and the log proves it.

    Evaluating an AI Pitch for Your Terminal?

    Book a free 60-minute discovery call with Frenchy Digital — a senior-led Black-owned LA agency. You leave with an integration assessment, a measured before-state for one workflow, and a fixed-price phased proposal within 5 business days. Call +1 (424) 272-5601.

    Evaluating an AI Pitch for Your Terminal?

    Book a free 60-minute discovery call. You leave with an integration assessment, a measured before-state for one workflow, and a fixed-price phased proposal within 5 business days.

    1517 S Bentley Ave Unit 204, Los Angeles CA 90025

    Frequently Asked Questions

    Sources & References

    Chris Machetto - CEO & Founder of Frenchy Digital

    Chris Machetto

    CEO & Founder of Frenchy Digital. Building apps and digital products since 2019 for startups and enterprises across LA, San Francisco, Paris, Geneva, and more globally.