Back to Writing
    Essay

    Reliable by Design

    29 September 2026 · Ryan Kerstein

    The tool was never the problem

    Adapted from a talk given to the 3rd International Quality and Patient Safety Conference, hosted by Dr Sulaiman Al Habib Medical Group, Riyadh, September 2026.

    In 2016 I was a surgical registrar, and most mornings began with the same conversation. A patient would arrive on the ward before seven, starved since midnight, unsure where to go and frightened about the operation ahead. Then they would wait, often all day, with no idea when they would be called. In theatre, the phone would ring mid-case: "When is Mrs Jones coming down?" And some afternoons, at five o'clock, that same patient would be cancelled and sent home hungry.

    Nobody designed that journey. It grew. When I mapped it, I found six handoffs with a gap after almost every one: told you need an operation, weeks of silence, starved from midnight, waiting with no time given, cancelled at five, home and no better. Every step had someone doing their job, and nobody owned the whole.

    So, like any surgeon who sees a pathology, I tried to fix it. Four of us built Lister, a patient communication and theatre scheduling platform. It gave patients the date, time and place of their operation from day one, sent nil-by-mouth prompts timed to the list, showed their live position in the running order and flagged every delay as it happened. We won the Hacking Surgery hackathon in London that year, and were finalists in the London Business School Health Tech Challenge the next.

    It worked. In a pilot in day case hand trauma at the Royal Free Hospital, 85% of patients using Lister felt well informed on the day of surgery, against 65% with standard written information. The share rating their experience "very good" doubled, from 40% to 80%.

    And then it died.

    Half the truth

    For years I blamed adoption. Health systems struggle to take up good ideas, and pilots without procurement budgets are theatre. I have given whole talks on this, and I still believe every word.

    The other half took me ten years to see. Lister changed one part of a system that had never been designed to keep a patient informed. Keeping the patient informed was nobody's job, so there was nothing for the app to be adopted into. The running order of the day did not bend when a patient knew more. The people around me had their own priorities and their own workarounds, and nothing in how their work was set up put patient information near the top.

    My app was never a bad idea. It was one fifth of a redesign.

    Why knowing has not been enough

    Most clinicians could sketch James Reason's Swiss cheese model from memory (BMJ, 2000). I have taught it for years, and the holes have not moved.

    Part of the reason is what we do when care fails. In my hospital, incidents are reported on Datix, and the first field on the form is the name of the person reporting. The investigation rarely gets much further than that box: reflect, perhaps retrain, move on. A second victim sits in the coffee room the next morning carrying the event alone, and the system that produced it waits for the next name.

    Elaine Bromiley shows where that leads. On 29 March 2005 she was a fit 37-year-old admitted for routine sinus surgery. At induction her airway could not be secured: the cannot intubate, cannot oxygenate scenario that every anaesthetist trains for. Two consultant anaesthetists and a consultant ENT surgeon kept trying to intubate for over twenty minutes. Two nurses in the room already had the answer. One fetched the tracheostomy set and told the team it was ready. The other booked an intensive care bed. Neither was heard. Elaine died thirteen days later.

    Everyone in that room was skilled, senior and trying their best. The knowledge was there; it sat with the quietest voices, and the room had no way to get it to the hands that needed it. Her husband Martin, an airline pilot, asked for learning rather than blame. His question built the UK human factors movement.

    The holes in the cheese are decisions, and most of them are made away from the bedside:

    • The rota was settled months earlier, in a spreadsheet, by someone who does not work on that ward.
    • The kit you reach for at 3am was chosen by procurement, often on price.
    • The pathway was written by a committee meeting at two in the afternoon, and is used at three in the morning.
    • Permission to speak was never written down at all. A culture decided it.

    Our procedures describe work as imagined: a full team, the kit where it should be, one patient at a time. Work as done looks different. At three in the morning you are on a ward you do not know, two patients are deteriorating at once, and the drug key has gone on break in a colleague's pocket. So you improvise. You borrow a key from next door, it works, and nothing is recorded. Most days that improvisation is invisible. On the worst day it fails, and it gets renamed error.

    The sociologist Diane Vaughan called this the normalisation of deviance. Before Challenger, NASA had seen damaged O-rings on flight after flight, and because nothing happened, the damage came to feel like routine. Drift never announces itself. Last year colleagues and I wrote about how NASA rebuilt its safety culture after Apollo 1, Challenger and Columbia, and what surgery can take from it (Annals of the Royal College of Surgeons of England, 2025). The uncomfortable conclusion is that our theatres drift in the same way their launch pads did.

    You cannot try harder

    Reliability science, much of it from the Institute for Healthcare Improvement, describes three levels. I ask every room to place its own unit honestly, on its busiest day of the week.

    1. Remind. Vigilance, training, posters and good intentions. About one failure in ten gets through, because attention decays. Train someone on Monday and by Friday the old habits are back; leave a poster up long enough and it becomes wallpaper.
    2. Standardise. Checklists, bundles and order sets, designed and measured. About one failure in a hundred.
    3. Design out. Systems that catch and correct their own failures before a tired human has to. About one in a thousand.

    Most of healthcare runs on the first level. Nobody works harder than clinicians; effort is not what is missing. You cannot try harder your way to one in a hundred.

    The WHO Surgical Safety Checklist is the best-known level-two control. In eight hospitals worldwide it cut deaths after surgery from 1.5% to 0.8%, and major complications from 11% to 7% (New England Journal of Medicine, 2009). Its power was never the paper. It was the pause, and the permission to speak. The first item is introductions, which is why in my theatres I introduce myself as Ryan: the most junior person in the room has to feel able to say stop. Where teams used the checklist as a structured conversation, it saved lives. Where it became a tick-box ritual, the effect evaporated.

    Aviation learned this earlier. When we interviewed airline pilots about their checklists, the difference was engagement rather than paperwork (Annals of the Royal College of Surgeons of England, 2022). Every item has a defined owner: one pilot calls it out and the other verifies it. For some checks, crews now read out the value they are actually looking at rather than simply answering "checked", a change one pilot described being introduced after an incident, to make the check more rigorous. And every flight ends with a debrief: what went well, what could be better, and why each pilot did what they did. Compare that with theatre, where one study found team members absent for the checks in 40% of cases and failing to pause in 70%.

    Level three looks different again. Since the 1950s, the Pin Index Safety System has kept gas cylinders off the wrong yoke: the pins sit in different positions for each gas, so a nitrous oxide cylinder physically cannot connect to the oxygen line. Nobody has to remember. The yoke remembers for them. That is what finished looks like: the easy action is the safe one, and the tired team at 3am gets the same result as the fresh one at nine.

    Five parts, not one

    Care is produced by five parts working on each other: people, tasks, tools and technology, environment and organisation. The model comes from SEIPS, the Systems Engineering Initiative for Patient Safety. Outcomes are what the five parts do to each other, so a change to one part is a change to all five, whether you designed it or not.

    Original presentation slide: Lister changed only tools and technology, while people, tasks, environment and organisation remained untouched

    Lister changed the tools. The tasks, the people, the environment and the organisation carried on exactly as before.

    This gives a test I now apply to every business case, and give to every room I teach. Before you approve a new tool or pathway, ask for the plan for the other four parts:

    • People: who does the extra work, and what do they stop doing to make room?
    • Tasks: which steps change, and which workarounds does it break?
    • Environment: does it still work at 3am, on a ward where nobody knows your name?
    • Organisation: who owns the outcome, and who pays once the pilot ends?
    • Tools and technology: what does it actually do, and what new failure does it bring with it?

    If there is no plan for the other four, you are approving a bolt-on.

    Human on the hook

    Healthcare does not lack innovation. We have hackathons, accelerators, pilots, awards and plenty of evidence. What we lack is adoption, and the five-part test explains most of the gap. It also explains what worries me about the current wave of AI, because every new technology closes one hole and can open another.

    • Diagnostic AI closes the missed or delayed diagnosis. It can open automation bias and deskilling: the reader stops looking where the machine says there is nothing, and the skill that would have caught its miss thins out.
    • Robotic platforms close a precision hole. They can open a skills-decay hole: surgeons who cannot convert to open when the platform fails, and teams who have never rehearsed it.
    • Ambient AI documentation writes my clinic letter while I talk to the patient. The letter still goes out under my signature, so verification has to be a real step, with real time in the clinic template and the job plan.
    • Alerts and dashboards close a visibility hole. Add enough of them and they open an alarm-fatigue hole.

    Which way each one goes is decided by the system design, not the product.

    The phrase I have come to distrust most is "human in the loop". Humans are poor passive monitors, and we defer to a machine most readily when it is confidently wrong. Lisanne Bainbridge described the irony of automation in 1983: automation keeps the routine work and leaves the human the rare, catastrophic catch, which is exactly the task we are worst equipped to do (Automatica, 198390046-8)). When the system fails, the inquiry finds a person conveniently positioned inside the loop. The clinician becomes the crumple zone, built into the design to absorb the impact of its failures. It is the person approach again, rebuilt around a machine, and it is still the clinician's name that goes on the Datix form.

    In 2016 I bolted technology onto people, and it died. Much of today's human-in-the-loop design bolts people onto technology, and it fails the same way. The difference is that today's vendors have the funding and the political weight to get their tools deployed regardless. Clinician-first design starts from what the machine does to the work, and only then asks who should watch it.

    Ten years apart

    The method, applied, is our skin cancer programme in Buckinghamshire. Urgent skin cancer referrals were rising year on year. Most turned out to be benign, and every one needed a specialist to say so. We could not hire dermatologists at the rate referrals grew, so patients with real cancers were queuing behind benign lesions.

    We wrote the problem brief first, with a named owner (me) and a definition of success: cancers found sooner, benign lesions out of the specialist queue, and no patient discharged without a safety net. Then we designed backwards from adoption. What does the pathway look like when this is simply how we work? Which named clinician runs it, and with what time in their job plan? What will the business case need to see? Only then did we ask what the pilot had to produce. Most organisations design the pilot first; it should be the last thing you design.

    Then we changed all five parts. Imaging staff were trained, with dermatologists providing oversight and taking the complex cases. Triage was redesigned so that benign lesions are discharged with safety netting. DERM, Skin Analytics' AI medical device, reads the dermoscopic images. Imaging moved into community hubs, out of the specialist clinic. Regional governance and audit were set up, and adoption was funded before we scaled. The programme is now business as usual, and it has cut our skin cancer referrals by 30%, safely. The tool was the smallest change we made.

    Original presentation slide comparing Lister in 2016 with DERM in 2026: Lister changed one of five parts, while DERM redesigned the pathway with ownership and adoption funding

    Lister had no owner beyond the founders, no adoption budget and an unchanged pathway. DERM had a named clinical owner with time, adoption funded before scale and a pathway redesigned around it.

    Designers, not passengers

    None of this is new to clinicians. When we take up a new operation, we watch it done before we do it, do our first cases with an expert present, audit our own outcomes honestly, open them to our peers at morbidity and mortality meetings, and expand only when the results hold. That discipline is reliability engineering by another name. Every surgeon has been a systems engineer since their first supervised list.

    Five moves turn a pilot into a reliable pathway:

    1. Name the owner. One clinician owns the adoption decision before the pilot starts.
    2. Write the problem brief. The hole, on one page, with a measure of success.
    3. Design all five parts. Tasks, people, environment and organisation. Then the tool.
    4. Fund adoption first. If the budget for afterwards does not exist, do not start.
    5. Audit, and stop when it fails. Measure honestly. A pilot that fails should end, in public.

    None of this is novel. All of it is unevenly distributed.

    If you are senior enough to sign a business case, you are one of the system's designers. The rotas, the kit, the pathways and the escalation cultures that cut the holes are all set at tables where people like you sit. A junior colleague can flag the system. You can change it.

    I did not become a better innovator between Lister and DERM. I learned what the other four parts were for.

    Reliability is designed, which means it can be designed by you. So pick one process where failure is easy, and redesign it until the safe thing is the easy thing. What is the one thing you would make reliable?


    Founders and investors sometimes ask me to pressure-test exactly these assumptions. How I work explains where I can and cannot help.