skip navigation
skip mega-menu

The Last Mile of Enterprise AI: Designing for Noise, Distance and No Signal

The Last Mile of Enterprise AI: Designing for Noise, Distance and No Signal

There is a version of every AI product that works beautifully. It runs in a quiet room, on a good laptop, over strong wifi, with two people speaking clearly in turn. It is the version in the demonstration, and it is the version the business case is written against.

Then it goes to a home visit. There is a television on, a dog, a partner interjecting from the kitchen, a phone that has been in a coat pocket all morning, and no usable signal. The professional is holding a difficult conversation and has roughly no attention to spare for software.

That gap between the demonstrated version and the deployed one is where most frontline AI programmes are actually decided.

What actually breaks at the frontline?

Not the model. Almost always, the conditions around it.

Multiple speakers and overlap. Real conversations interrupt, overlap and trail off. Two people finishing each other's sentences is normal human speech and a genuinely hard engineering problem. Attribution errors matter far more than transcription errors here - a sentence assigned to the wrong speaker does not read as a mistake, it reads as a fact.

Accents, dialects and speech differences. Performance varies across accents, and across speech affected by illness, medication, distress, dentures or fatigue. In services working with older people and people with disabilities, the speech that is hardest to process is often the speech that matters most. This is an equity issue as much as an engineering one.

Telephony. A large share of frontline conversation happens on the phone, where audio is compressed, bandwidth is narrow and quality is inconsistent. A tool tested only on in-person or online meetings has not been tested on a major part of the job.

Background noise. Not white noise - structured noise. Television, radio, other conversations, traffic, machinery. Speech-shaped interference is considerably harder to handle than a constant hum.

Connectivity. Rural areas, basements, lift shafts, thick-walled buildings, stairwells. Anything that requires a live connection to work is unavailable somewhere in the patch, and it is rarely the same somewhere twice.

Devices. Personal phones, shared devices, ageing tablets, whatever was in the pool cupboard that morning. Battery is a real constraint: a tool that drains a phone by lunchtime is a tool that will be closed at eleven.

Attention. The one nobody specifies. In a difficult conversation the practitioner has none to give. Anything requiring more than a single deliberate action - start, stop - is competing with the person in front of them, and it will lose.

Why do demonstrations hide all of this?

Not through dishonesty. Through selection.

Demonstrations are run in the conditions where the product works, by people who know how to use it, on prepared audio. Pilots are run by volunteers who want it to succeed, who tolerate friction a sceptic would not, and who often self-select the easier conversations to try it on. Both are useful. Neither tells you what happens when the tool reaches a practitioner who did not ask for it, on a Tuesday, in a house with a television on.

This is also why headline accuracy figures should be treated with caution. An accuracy percentage without stated conditions - how many speakers, what noise environment, what audio source, what accents - is not a measurement. It is a marketing artefact.

What should the tool do when it fails?

This is the question that separates products designed for the field from products adapted to it, and it barely appears in requirements documents.

Failure at the frontline is not an edge case, it is a weekly event. What matters is the behaviour around it:

  • Never lose the capture silently. If processing fails, the audio should still be there. Losing a conversation a practitioner cannot repeat is unforgivable, and one occurrence will end that person's use of the tool.
  • Degrade rather than stop. No connection should mean capture now, process later - not a blocked screen and a wasted visit.
  • Recover from interruption. Calls drop, batteries die, apps get backgrounded by an incoming call. A resumed recording should reconcile into one coherent record, not three fragments.
  • Be honest about uncertainty. Where the system knows a passage was poorly captured, say so in the output. A flagged gap is a manageable problem; a fluent summary papering over an inaudible two minutes is a dangerous one.
  • Stop instantly. One movement, no menus. This is a dignity requirement as much as a usability one.

How do you test for any of this before buying?

Not with a demonstration. With a field trial designed to be unkind.

A workable protocol:

  1. Sample your actual conditions. Take a fortnight and record - with proper consent and governance in place - the range of real environments: in-person visits, phone calls, noisy settings, rural locations, the specific accents and speech patterns of your population.
  1. Test against those, not against clean audio. Any supplier confident in their product will accept this. Reluctance is itself information.
  1. Measure what matters. Not overall word accuracy - speaker attribution, and whether critical information survives into the summary. Have the practitioners who were in the room mark what should have been captured.
  1. Break it deliberately. Kill the connection mid-conversation. Let the battery die. Take a call halfway through. Watch what happens to the recording and what the user has to do next.
  1. Test with sceptics. Include people who did not volunteer. Their tolerance for friction is the tolerance you will actually be deploying into.
  1. Test on the devices you actually own, including the oldest ones still in circulation.

The point is not to catch a supplier out. It is to know where the boundaries are before you write them into a contract, so that adoption problems in month three are anticipated rather than surprising.

Why this decides adoption, not just performance

Frontline tools get roughly two chances. A practitioner who loses a conversation, or who spends longer fixing a summary than writing the note would have taken, will revert to the old method - and they will tell their team. That informal verdict spreads faster than any training programme and is close to impossible to reverse.

Which means field reliability is not a technical requirement competing with the interesting features. It is the requirement that determines whether any of the other features are ever used. A tool that is slightly less capable but reliably works in the car park will beat a more capable tool that works in the office, every time, because only one of them is being used at four o'clock on a wet Thursday.

The organisations that get this right test for the worst day rather than the best demonstration. It is a less impressive procurement exercise and a considerably more successful deployment.

At VE3 we work with organisations on the integration, resilience and deployment design behind AI that has to work outside the office. If you are planning a field trial, we are happy to talk it through. Contact Us.

Subscribe to our newsletter

Sign up here