Nobody Wrote This Down

Seven years of the MicroBinfie podcast

Nabil-Fareed Alikhan

Lee Katz

Andrew J. Page

30 September 2026

Seven years of the MicroBinfie podcast

Copyright © 2026 Nabil-Fareed Alikhan, Lee Katz and Andrew J. Page

The right of Nabil-Fareed Alikhan, Lee Katz and Andrew J. Page to be identified as the authors of this work has been asserted by them in accordance with the Copyright, Designs and Patents Act 1988.

All rights reserved. No part of this publication may be reproduced, stored in a retrieval system or transmitted in any form or by any means without the prior written permission of the authors, except for brief quotations in a review or scholarly work.

This book is built from the MicroBinfie podcast, recorded between 2019 and 2026 and available at soundcloud.com/microbinfie, with episode notes at microbinfie.github.io. Quoted speech is transcribed from those recordings and reproduced verbatim at the word level; punctuation within quotations is editorial, since the recordings do not contain any. Material contributed by guests remains theirs, and is used with permission — see Thanks for the people this book is made of. Guests, and the colleagues they talk about, appear under pseudonyms: a name marked with an asterisk (*) at its first appearance in a chapter has been changed.

Every effort has been made to check the factual claims in this book against published sources. Where a recollection on air turned out to be mistaken, the correction is in the text and the original is recorded in the project’s provenance files. Where a prediction turned out to be wrong, the prediction stands as it was made. Any errors that remain are the authors’.

The views expressed are those of the authors and of the guests speaking for themselves, and are not those of any employer, funder or institution named in the text.

First edition, 2026

Published by Nabil-Fareed Alikhan, Lee Katz and Andrew J. Page

Typeset from Markdown. Set in Georgia.

Contents

Foreword

  1. Things That Went Wrong

  2. Software Not To Write

  3. What Language Should I Learn

  4. The Formats Nobody Designed

  5. How This Field Got Here

  6. Early Days of MLST

  7. Assembly, and Other Acts of Faith

  8. Sequencing Technologies, and the Arguments About Them

  9. Trees, and How to Argue About Them

  10. Bayesian Magic

  11. Sketches, Hashes and k-mers

  12. What Is a Species Anyway

  13. Naming Things

  14. Pangenomes

  15. Mobile Genetic Elements

  16. Resistance

  17. Bugs With Personalities

  18. Outbreaks

  19. When It Was Not a Drill

  20. The Pipeline Wars

  21. Containers, and Other Ways to Stop Suffering

  22. Ontologies, and Other Things Nobody Thanks You For

  23. Who Gets to Sequence

  24. Publishing, Reviewing, Shouting Into the Void

  25. Machines That Write Code

  26. What Happens When the Money Stops

  27. Teaching the Thing Nobody Was Taught

  28. Careers Nobody Planned

  29. Genomics at the Cinema

Thanks

A Note on Sources

Glossary

About the Authors

Colophon

Index

Foreword

There is no manual.

That is the line the podcast opened with, week after week, for seven years, and it is the reason any of this exists. The full version runs: there is so much information we all know from working in the field, but nobody really writes it down, there is no manual, and it is assumed you will pick it up.

Everyone in this field has picked something up that way. You learn that a file of nothing but capital letters in the quality string is old data lying to you, and you learn it because somebody mentions it in a corridor, not because you read it anywhere. You learn which assembler to reach for, which conference bar has the useful conversations, that the tool everybody cites has not been updated in four years, and that the person who wrote it knows and feels bad about it. None of that is in a paper. Some of it is in a thread that has since been deleted.

The corridor is the problem, because most bioinformaticians do not have one.

An organisation that needs bioinformatics usually needs exactly one bioinformatician. You are the expert, which sounds better than it is. It means there is nobody in the building to ask, nobody to mention the thing everybody else already knows, and nobody to tell you your idea is bad before you have spent three months on it. The work is specialist enough that most places never need a second person doing it, which is how a whole field ended up as several thousand people working alone, in different buildings, solving the same problems separately and never finding out.

The three of us met the way most people in this job do: at conferences and hackathons, which are the same thing with more laptops and better arguments. What we wanted afterwards was somewhere to keep having those conversations — a safe space to share what we knew with peers, rather than with an audience we had to impress.

The first attempt was a virtual lab talk series, coordinated through Slack. The seminar you would get down the corridor if you had a corridor, put online so it did not matter where anybody was, and posted to YouTube. It fizzled out. The reason is obvious in hindsight and was not obvious at the time: a lab talk has to be prepared, and the people we were asking were academics and scientists who were already overworked. Volunteers were hard to come by. Almost nobody actually said no. They just never had a free evening.

So the three of us — Nabil-Fareed, Lee and Andrew — did a podcast instead, mostly for ourselves. Three people talking is easier to arrange than one person presenting, and it needs no slides. If anybody else listened, that was a bonus.

Enough of them did. A hundred and fifty-six episodes, seventy-three hours, and a back catalogue that nobody, including us, has ever listened to end to end. Along the way a great many people agreed to come and talk to us, always volunteering, and often at a conference. They are named at the back, and the good bits are mostly theirs.

In the chapters themselves, though, they are not. The guests, and the colleagues they talk about, appear under changed names, each marked with an asterisk the first time it turns up in a chapter. The episodes are public, and everybody on them chose to come on and talk. Even so, we felt that the recent political assault on public health warranted blanket anonymisation, for the welfare of all of those involved: a name in a printed book lasts longer, and travels further, than one in a podcast feed. The words are still theirs, and so is the credit.

This book is that catalogue, rearranged into something you can read.


Three things worth knowing before you start.

We have not tidied up our own predictions. Where somebody said a thing on air that the following years demolished, the wrong version is still here, and the correction follows it. There is a chapter in which one of us explains in 2023 that hiring will have to change, because reading code carefully is the skill that matters — and, three years later, on the same podcast, that wanting to read code line by line is an immediate no. Both are in. Going back and quietly winning old arguments is the one thing this book must never do.

We have not resolved the disagreements either. There is a three-part argument about workflow managers in here that does not conclude, because it did not conclude. Lee still thinks make is a perfectly good workflow manager, and says so, and everyone still shoots him down. The same is true of Perl, and of about four other things.

Some of this has dated, and that is the interesting part. Sixty-five of the hundred and fifty-six episodes were recorded in 2020 and 2021, and most of that was reporting rather than thinking. What survives is the small number of conversations about what it is like to be asked, on a Tuesday, to build a national capability by Friday. The rest of it — the round-ups, the weekly variant updates — is a different book, and it should be.


You do not need to read this in order. The chapters are grouped by subject and most of them stand alone, which is a property of the source material rather than a design decision: each episode was made to be listened to by somebody who had not heard the others.

But if you want a route: the file formats and the things that went wrong are the easiest way in, the taxonomy chapter is the argument that runs longest and develops most, and the last chapter is three bioinformaticians watching bioinformatics-heavy movies and being completely unable to stop commenting.

Nobody wrote this down. So we did.

Things That Went Wrong

In the first years of sequencing Listeria in real time, the numbers started moving in the most exciting direction available. There was a cluster running across several states, bigger than anything anyone had seen before, and it was still growing.

It was the sheep. The Listeria isolates were being grown up on blood agar, “there was a Listeria outbreak in sheep and we had no idea”, and so a little of the thing being hunted arrived already sitting on the plate. The largest Listeria cluster on record was the medium. It got written up eventually under a title with very little fanfare: “two Listeria monocytogenes pseudo outbreaks caused by contaminated laboratory culture and media”1.

The same shape turns up outside our field and at much worse stakes. German police spent sixteen years chasing a woman whose DNA was at scene after scene, because “they just kept seeing the same signature over and over again for all these murders all over the country”, and she worked at a factory in Austria making the cotton swabs the police were collecting the evidence with. The version we tell each other has her licking her fingers.

Neither of those is a sequencing failure. Both systems did exactly what they were told, and the DNA they found was really there. What was wrong was where it had come from — Listeria arriving with the culture medium, human DNA arriving on the swabs — and in each case that had happened earlier and somewhere else, before anybody had looked at a result. Nothing in this work breaks where you are looking at it.

Too good to be true

“Is what you are doing worth the Nobel Prize, or is it contaminated?”

That is the question, and the awkward part is who it has to be asked of. Not the paper you are reviewing — your own good news, at the moment you are least inclined to hear it, plot already on the screen and everybody in the room pleased. So it gets pushed earlier, into a set of checks you run before you are allowed to have a feeling about anything.

Those checks begin by not taking the microbiologist’s word for what is in the tube. Asked whether that was really the position — that you would not simply believe somebody who told you this one is E. coli and that one is Salmonella — the answer came back without a pause.

“No, never assume anything, it makes an ass of you and me.”

The first pass is a database: Kraken, or Mash to find the closest relative in RefSeq, or the Genome Taxonomy Database, which was the new arrival at the time and has stayed. What makes it more than a lookup is the framing — you hold the single isolate as a null hypothesis and let conflicting taxa disprove it. Which is fine until you remember that a database is somebody else’s data too. Lee built Kalamari, a curated set for exactly this job2, by going to subject-matter experts one at a time and asking which genomes they expected and which were the usual contaminants. He selected completed genomes, or ones as close to complete as possible, because fragmented public assemblies carry passengers of their own. A contamination screen that is itself contaminated is a long week.

Catching a thing you cannot name

Databases only find what somebody has already put in them, which is why the more useful checks never ask what the contaminant is. They ask whether the data is internally consistent with having come from one organism.

Start with shape. A Salmonella assembly from current Illumina should land around 4.5 megabases in a couple of hundred contigs, so a twenty-megabase Salmonella, or one in two thousand pieces, is a problem you have identified without identifying anything. Then base composition: plot GC per sequence and two organisms with different GC separate, giving you two clouds where one organism would have given one.

Then the one worth learning properly, which is the k-mer count histogram. Chop the reads into k-mers, count how many times each k-mer occurs, and then — the step that trips people the first time they meet it — count how many distinct k-mers occurred that many times. Counts of counts. A clean single-genome library gives one hump sitting at roughly your coverage, with a spike at the low end that is sequencing error. Two organisms present at different abundance give two humps, because they were sequenced to two different depths, and the same trick shows human reads sitting in among microbial ones. What it cannot do is tell you which k-mers belong to which thing, so it is never a verdict, only a reason to go and look properly. The real argument for it is that it runs in about a minute, and a check nobody has time for is not a check.

After that, single-copy genes. Something you have exactly one copy of should come back as exactly one copy, and when it doesn’t, the first suspect is a second genome. BUSCO carries a panel that works across the tree of life, CheckM does the equivalent job for bacteria, and for a species you already know you can run seven-gene multi-locus sequence typing (MLST) and simply count: seven loci, seven alleles. A mixed allele call at a locus means two organisms are disagreeing about what belongs there. MLST gets declared dead roughly once a year and keeps quietly working as a quality metric, which is a better retirement than most methods manage.

None of it replaces looking. Map the reads to a reference, open the result in IGV or Artemis, and “just by eye you can see things that a computer can never tell you”. Coverage dropouts are the standard example, because they slip past everything else — the Q-scores are fine, the read count is fine, and a large stretch of the genome is simply absent because something happened in the library. Miss that and you can spend a fortnight explaining a deletion that was never there.

Coverage is worth reading even when nothing has dropped out of it, and here the tip is the kind that is invisible until somebody says it out loud and obvious forever afterwards: “the coverage across a genome should never be constant”. It should fall away from the origin of replication towards the far side, because in a growing culture the cells are partway through copying themselves and there are simply more copies of the sequence nearest the origin. What a flat profile doesn’t do is prove anything. How steep the slope is depends on how fast the cells were dividing when the DNA was extracted, and a culture that had already gone into stationary phase can come back nearly flat, so flatness is a question about how the sample was grown rather than a verdict on it. Coverage that instead tracks GC is a library bias that has no business being there. The name arrived at for this skill, after two failed attempts at describing it on air, was that “you’ve got to be contamination sensitive or something”.

Several of those artefacts make more sense once you remember that “all of the work we do is basically we’re just taking pictures” — an Illumina run is a camera photographing a slide, and its failures are photographic. Put too many clusters down and their signals smear into each other, and the quality can be genuinely poor while the report calls it excellent. The camera is not the exotic device the specification sheets imply, either; “it’s probably just a little bit better than your iPhone camera”, which is why the fix was fewer clusters rather than better optics. And sometimes the chemistry simply stops, and a cycle arrives with A’s and T’s and G’s in it and no C’s at all, which should not be possible. “I think we’re at the bleeding edge of science and sometimes things don’t work.”

The one that catches everybody is the barcodes. Index sequences are short, six to nine bases. What matters is the Hamming distance between the indexes you happened to order — how many errors a barcode can absorb before it reads as one of the others — and the standard demultiplexer will forgive about one. The wet-lab word for the result is bleed-through, but this version of it isn’t chemistry at all: “it’s really a hash collision in your barcode”.

There’s a chemical version too. The classic stories are of indexes arriving from the supplier already slightly contaminated, and once they’re dirty on the way in, all you can do is spot it and deal with it. Dual indexing buys a good deal more tolerance, at the cost of sequencing you would rather have spent on the genome, and doing the demultiplexing yourself gets you a report that flags barcode trouble. We insist on twenty bases for a primer to be specific enough, then hang an entire run off eight.

Some of it did not happen in your experiment at all. Data comes back carrying the signal of something alarming, and the alarming thing was on the machine two days earlier or growing on the bench beside it. One mechanism is a liquid-handling robot with fixed tips - metal prongs rather than a box of disposables, dipped into every sample in turn and rinsed with water in between, because anything strong enough to sterilise them would take them apart. Which makes it a purchasing decision rather than a protocol: “why are we buying tens of thousands of tips when we could just have this one fixed tip robot”. After PCR, when all you are doing is moving DNA about, that is the right answer. Before it, you are inoculating one sample from another to save on plastic.

The control you do not tell them about

All of that assumes the problem is in the data. Often it is a person, and there is no software for a person. Contamination that arrives by hand — somebody introducing something awful into your data set — is the category with the worst outcomes and the fewest defences.

A thousand-sample project, first few plates clean, then sample after sample coming back as a mixture of strains. Somebody in the lab had worked out that picking single colonies and culturing them up was an afternoon that could be saved, on the reasoning that “there’s one bead it’s probably got one bug on it”, and had started sequencing straight off the bead. It is not an unreasonable thought, and the person who had it was trying to save everybody time. The only thing wrong with it was “but they didn’t tell anyone”, and it cost weeks of work, a great deal of sequencing, and the whole set done again from scratch.

Which is what controls are for, and why the good ones are dishonest. A blank through the extraction. A positive control on the plate, and if you are using an external provider, “you should always put on a control randomly on the plate and don’t tell them”. Better still is deliberately mixing sequence types or serovars across a plate, because a plate rotated in the lab is undetectable when every well holds the same organism and obvious the instant one well is a different species from its neighbours. In low-biomass work the controls stop being a formality altogether — there are studies where the blank came back carrying more taxa than the sample did, which is how “you end up with a placenta microbiome”. That one got settled in public: a study large enough to carry its own controls went looking and found no placental microbiome at all, only the reagents3. The signal had been the kit. And run cases and controls together, same week, same reagents, same technician, because a batch effect is “one small thing that just permeates your entire data set” and it will walk you all the way down the garden path before you notice. Controls belong in a study as first-class members of it, not as something remembered in the fortnight before submission.

Which language should I learn

A different kind of thing goes wrong, and it is not in the sample.

Ask a room whether to learn Perl or Python and you do not get an answer, you get a position. Perl is very good at text munging and turns messy if you are not disciplined about it; Python is prettier and has roughly a couple of ways to do a thing where Perl has a million. The field moved across in the space of a few years, and the number everybody points at is the daily uploads — twenty or thirty to Perl’s CPAN archive against seven or eight hundred to Python’s PyPI — which reads either as a language that has matured or as a language everyone has left, depending entirely on which of the two you write in.

“It’s dead, and I keep programming in it, so it must be alive.”

The replacement did not help. “Perl 6 has been in development hell for, what, 14 years now” was the verdict, and the eventual resolution was to stop pretending it was a successor at all, rename it Raku, and let the two be separate languages.

None of that is the hazard, though. The hazard is inheriting a script where “you might have 500 lines with no comments and it’s all just one big blob”, written by a student who submitted and left, in whichever language happened to be in fashion that year. “Bad code is bad code”, and no syntax anybody has designed has ever prevented it. The argument about languages is a maintenance problem wearing a costume.

Where it stands

Somebody once turned up wanting to learn bioinformatics from a standing start and leave with a Nature paper, and had set aside a week for it. Asked on air how long that actually takes, the one of us they had visited had a figure ready:

“Approximately four and a half days.”

What a week can honestly do is give somebody the flavour of the thing. Callum Devlin*, who was teaching graduate courses at a Canadian university when we spoke in 2020, would add the vocabulary: to his mind the core of a short course is teaching people to help themselves, with the terms, the man pages and the help flags, because without that very little of an intensive week is retained.

The rest is the room. There is always one person at the back who could finish every exercise in two seconds, and in the same row somebody who is not sure where the tab key is. The two things Callum found missing most often are not commands at all but the file system and the idea of a function. People are scared off by the bare cursor, which is not cowardice; a prompt tells you nothing whatsoever about what it will accept.

The response has been abstraction, and it works. Point-and-click workflow systems let wet-lab colleagues assemble, annotate, BLAST and get a result without knowing how an assembler works, and for a well-tooled organism that’s a good outcome. The bacterial path is paved, and Callum’s point was how unusual that is. His PhD was in the microbial eukaryotes, doing single-cell transcriptomics on Paramecium bursaria, a cell that carries two nuclei, a population of algal endosymbionts and a giant virus loitering on the outside. There nothing is routine, the tooling needs tweak after tweak, and annotation generally means custom-trained models rather than just running Prokka. We’d seen the same gap from the other side, with eukaryotic genomes dragged out one at a time while the bacteria next door came through by the thousand.

Whether the paving is a kindness split the conversation. Callum’s position was that learning a workflow system well takes real effort, and that depending on what somebody wants to do next, the effort might be better spent on the command line. The reply from our side was that it doesn’t, and Galaxy is easy.

The cost Callum worried about isn’t a wrong assembly. It’s that somebody can drop a sequence into BLAST and announce that SARS-CoV-2 was built in a laboratory, and that with everything abstracted away “you can’t even begin to ask those questions”. Even the one person in the room who wants to know what the sort step is doing can’t reach the flags. That was recorded in the autumn of 2020, when it was not a hypothetical, and nothing since has made it less true.

The other thing that has dated is a prediction. Asked in early 2020 whether the newer languages would displace Perl and Python, the answer was “maybe but we’ll see in about 10 years”, which leaves four years still on the clock. One of us had already jumped, on the grounds that leapfrogging was cheaper than catching up, and gone off to learn Rust when the bio library for it was not finished. That bet is looking better now than it did then.

None of these disasters was caught by a threshold. Every one was caught by a person who thought something looked odd and could not leave it alone, and what let them do it was a pile of small facts about how assemblies and cameras and lab plates actually behave. “Maybe it’s not even written down but you know it anyway.”

Notes

  1. Matanock A, Katz LS, Jackson KA, et al. (2016) Two Listeria monocytogenes pseudo-outbreaks caused by contaminated laboratory culture media. Journal of Clinical Microbiology 54(3):768–70. doi:10.1128/JCM.02035-15
  2. Katz LS, Griswold T, Lindsey RL, et al. (2025) Kalamari: a representative set of genomes of public health concern. Microbiology Resource Announcements 14(2):e00963-24. doi:10.1128/mra.00963-24
  3. de Goffau MC, Lager S, Sovio U, et al. (2019) Human placenta has no microbiome but can contain potential pathogens. Nature 572:329–34. doi:10.1038/s41586-019-1451-5

Software Not To Write

“I actually do have my own variant caller on GitHub. I haven’t deleted it, but I should out of just embarrassment.”

For about a decade everybody had one. Everybody had an aligner as well, and a wrapper round somebody else’s assembler, and a script that turned one tab-delimited file into a slightly different tab-delimited file. “Every PhD student poster could write their own variant caller on Python. Probably still do in some places.” It was never that the existing tools were bad. Writing a thing is more fun than reading about a thing, and reinventing the wheel is easier than going away and reading the papers that tell you not to.

So the first two conversations we ever recorded were an attempt to talk the field out of it, and they were, honestly, a literature review in disguise. The test we came up with had three parts. Do not write it if there is already a plethora of tools doing the job. Do not write it if the problem is more or less solved, or has been shown to be unsolvable. Do not write it if the technology underneath has gone.

Applied, that test is brutal. Multiple sequence alignment is about the hardest problem in the space and the two tools everyone actually reaches for were published in 2002 and 2004, so a new one has to beat twenty years of tuning before anybody will look at it. Short-read mapping is finished — BWA, Bowtie2, and half a dozen others, and the author of minimap2 has published a blog post politely telling you to use BWA instead if your reads are short. Variant calling is barely a program at all; it is a set of filters over a pileup, and whichever one you pick you will end up making the same handful of tweaks about coverage and allele fraction and how much you trust the ends of reads. Phylogenetics is worse again, because the people maintaining RAxML compile it against particular instruction sets to squeeze the last few percent out of your processor, and you are not going to out-engineer that on a Thursday.

There is a joke buried in the mapping half of that. BWA-MEM, which is what a great deal of the world’s short-read data is actually aligned with, was written up for publication and rejected on the grounds that it was no longer novel. It has sat on a preprint server ever since, “just pulling in vast numbers of citations”, and our verdict on the editor responsible was not a generous one. SMALT, from the Sanger Institute, was still unpublished when we discussed it in 2019. It was being used in plenty of papers regardless, but not having a paper of its own “puts a lot of people off” using it. Being the tool everybody reaches for and being publishable are two different things, and only one of them is what a career is scored on.

The obsolescence test is the one nobody applies to themselves. Around 2010 the big sequencing centres put their microarray equipment “in the skip. I don’t know what that is in the US”, and moved to RNA-seq. Papers using microarrays went on climbing for another four years, peaking in 2014, and were still being published in the tens of thousands in 2018. The instrument was in a skip and the literature had not been told. That is what it is to “get caught in this long tail of science” — not wrong exactly, just working steadily on the far side of a change that has already happened.

None of which convinces anybody. We know it does not, because it did not convince us either. The argument that works is not about whether the tool is needed. It is about what happens after you have written it.

What you are actually signing up for

Asked what to say to somebody about to start a new tool, Tristin Lindemann* — who has written more of the ones people actually use than anybody sensible would admit to — did not hedge.

“My first advice would be to don’t write a new tool.”

Not because the idea is bad — because it is a commitment, and because the only thing that reliably makes software good is being trapped with it. You will never make a tool useful and reliable “unless you are forced to use it and suffer through its problems”. Software written to be published and then left behind is software nobody has ever debugged while actually needing the answer.

The arithmetic is unkind. “The average lifespan of bioinformatics software therefore is about three to five years”, which is not a technical figure at all. It is the length of a grant, or of the contract of the person who wrote the thing. Velvet was the backbone of a great many pipelines and stopped being updated because the person developing it went off to do other wonderful things. A seven-gene multi-locus sequence typing (MLST) web server one of us built during a PhD went quietly offline five or ten years later, and by then there was nobody left who even knew who to ask about putting it back. Papers get published pointing at web services with no source code attached, and when the PhD or the postdoc project ends, that is it, it is gone.

Andrew changed institutions and left a trail of software behind for his old colleagues to field the bug reports on, and the honest position on that is not a noble one: “If I’ve published something five years ago, then should I really be maintaining it now for free?” Nobody has a good answer. Companies asked for support as well. He told them it had to be paid for, and “when they start hearing proper commercial rates they run away”.

Tristin described the other half of it: the guilt that comes with pull requests he didn’t have time to review and issues he didn’t have time to investigate, on software people depended on. He also counted himself lucky, because his employer saw maintaining it as an important part of his job, and he knew that wasn’t true for a lot of people.

The funding does not exist because maintenance is not interesting. It is not sexy, and funders want novelty and new technology, so the way you keep a tool alive is to disguise it as a new one — a version two, a rebrand, the same method moved from bacteria to yeast to viruses, or “they put in the cloud because cloud things are obviously so much fluffier”. There are a small number of exceptions where somebody decided a piece of software was infrastructure and funded it as infrastructure — samtools, BLAST, Galaxy, FastQC — and you use all of them without thinking about it, which is the point.

There is also a quieter problem, which is that most of what gets written should never have been released in the first place. Reviewing software submissions turns up a specific failure: people who know enough to produce a package but not enough to see that the package is a bespoke thing for their own project.

“That’s a polite way of saying it’ll only ever run on their laptop.”

Versions, and why they are load-bearing

Here is the plain version. Semantic versioning gives a piece of software three numbers. Change the first and you have broken something people depended on — the interface, the options, the input format. Change the second and you have added a feature without breaking anything. Change the third and you have fixed a bug, maybe no more than a typo. Anyone reading the three numbers can tell how frightened to be.

Here is the honest version. The three numbers are how you answer the worst email you will ever receive, which arrives about a year after publication and says: “Oh yeah, I ran your code, I’ve got this answer. How come your new code isn’t producing the same answer?” If you tagged the release, you can go back to the version they had, run it, reproduce their answer, and start looking for the change that moved it. If you did not, there is no answer available to you at any price, and the paper that used the tool now rests on a thing that cannot be reconstructed.

That failure scales. Trying to benchmark assemblers honestly turned out to be nearly impossible for exactly this reason: the versions moved faster than the comparison, so by the time you had finished, “whatever point you wanted to make, it’s gone, it’s obsolete”. A comparison of tools without version numbers is not a weak result, it is not a result. And it is not only the versions — it is the whole apparatus around them. Documentation for users rather than for developers. A worked toy example, because a toy example tells someone in ten seconds what an options list does not tell them in an hour. A licence, because open source is not a licence and somebody in a company genuinely cannot use your code until you have said which one it is. Tests, which are the thing that lets you change the software at all, and which are still taught as an afterthought: “I did a degree in software engineering and we had one little module on testing in four years.”

The rule of thumb offered was half your time on tests and half on code, which sounds absurd until you notice that the half spent on tests is the only half you can revisit safely.

Underneath all of this sits the difference between where the software was written and where it ends up. Research is allowed to be adventurous; you tweak protocols left, right and centre and you may not run controls every time. Public health doesn’t work like that.

Tristin moved from one to the other and found all the rules and steps frustrating at first, then saw they were there for good reasons. During the pandemic his lab’s genomic results uncovered a major hotel quarantine outbreak and got a residential tower block locked down, and in his words “these decisions, the decisions that get made from what we present, are real decisions”. Documented processes and fixed version numbers are what get software from one side to the other.

The bit where somebody has to pay

There is a price point for scientific software, and it is not a coincidence. Commercial tools get priced at something just below the level where you would think about hiring somebody for a year to do the job instead — a postdoc for a year costs you a hundred thousand, so the licence comes in at eighty, and the vendor knows you will not hire a developer to avoid it. The same logic is visible in miniature in a hardware quote one of us pushed back on, having explained what the actual budget was: “And then magically it was $1 less than the quote.”

You can be annoyed about that, and we were. But it is worth being clear about what the money is buying, because it is not the algorithm — the algorithm is usually in a paper you can read for nothing. What you are buying is that somebody will still be answering the phone in five years, that the version numbers mean something, that the validation exists. Beyond a certain point nobody in a clinic wants a tool at all: “They will just want a single button that you push and that they know it’s validated.” Academics avoid paying for software almost as a reflex, which is fine right up until the moment the thing has to work every time, and then someone, somewhere, is paying.

Where it stands

That advice went back out eighteen months later, with a global emergency in the gap, and the things it complained about had not moved an inch: the install-then-fail ritual, the missing URL in the abstract, the tool that only runs at the institution that wrote it.

What has moved is the field’s willingness to take the advice. The tools named in those first two conversations are, to a slightly depressing degree, still the tools. The saturation argument aged well.

The predictions did less well, and one is worth putting on the record rather than quietly forgetting.

“16S is dead. Please do not do any more tools for it.”

It was not dead. It is cheap, it is established, and people are still publishing amplicon studies now. The related prediction — that short reads were heading the way of microarrays - has aged into a maybe: long reads did become the interesting place to write software, exactly as claimed, but the skip is still empty.

And the strongest single counterexample to our own argument is annotation. Prokka1 was the tool in every pipeline, and by the terms of that test nobody should have written another one. Somebody did: Bakta arrived in 20212. Prokka had its weak spots, and one of them was already on the record. In 2019 Tristin told us he regretted the step that writes its GenBank files, which relies on a binary from the National Center for Biotechnology Information (NCBI) that expires every six months to a year, with no warning until it breaks.

A different kind of counterexample turned up in 2026. Rufus Georgiades* had already written one host-read removal tool, Hostile, which ran reads through minimap2 and Bowtie2, and he came to think the approach was “incredibly wasteful because you’re spending all this effort aligning human reads that you’re about to throw away” and hard to scale to human pangenomes. So he wrote Deacon, which stores minimisers instead. A better idea turned out to be reason enough for a new tool. What it doesn’t change is the other half of Tristin’s advice: whoever writes it has to be the one who’ll be there when it breaks.

Which is roughly where we came in. When we recorded all this, the variant caller was still up there, and taking it down would have meant going and looking at it.

Notes

  1. Seemann T. (2014) Prokka: rapid prokaryotic genome annotation. Bioinformatics 30(14):2068–9. doi:10.1093/bioinformatics/btu153
  2. Schwengers O, Jelonek L, Dieckmann MA, et al. (2021) Bakta: rapid and standardized annotation of bacterial genomes via alignment-free sequence identification. Microbial Genomics 7(11):000685. doi:10.1099/mgen.0.000685

What Language Should I Learn

“Learn Python, then R, then something like C, C++, C# and Java, and then you can learn the rest. So there you go. You can turn off the podcast now.”

Ninety seconds in and the question is answered. It had been put by a PhD student after a talk — what should I learn, and once I have learned that, what next — and it came back as an ordered list with no hedging anywhere in it, which is most of what makes it a good answer. It is also a route that nobody in the room had taken.

The person who gave it started on C++, then Java, then Matlab, Maple, Prolog and Caml as an undergraduate, then Perl, then PHP, then went out into a job writing Ruby, then back to Perl, then C, and arrived at Python somewhere near the end. Another of us picked up C in a college course, stopped using it, learned PHP to do web work — “which is also nauseating” — and met Perl in graduate school in 2004, because that was what graduate school taught. The third did a Java degree with Matlab and R and a bit of SQL sprinkled through it, then wrote a solid year of Perl on a first research project “because that was what the lab did”, and was allowed to move the work to Python only “after I was more familiar and could be trusted”.

Three careers, and not one of them contains the sequence recommended at the top of this chapter. What they contain instead is a list of options. The ideal path is ideal partly because nobody walks it, and none of us came anywhere near.

Whatever was already in the building

The competing answer arrives immediately and it is not a language: “look to who is in your space and what are they using and who’s willing to help you”. The reasoning is that the concepts are the part that transfers. Once you have them — a conditional, a loop, a primitive, and how those assemble into something that does a job — the next language is largely a matter of finding out where the semicolons go, give or take the ones built on a different paradigm entirely. What does not transfer is having somebody down the corridor who will look at your screen.

Which is also, read back, a description of what happened to all three of us. That year of Perl was not a decision about Perl. The reason given for it is that a new arrival does not walk in and start rearranging the furniture: “I’m not just going to come in and just say do something else.” In a vacuum, one of us would still send a beginner to Perl out of straightforward affection for it, and says so. But “we don’t work in a vacuum” — the work happens on shared machines and in shared repositories, and a language only you can read is not a preference, it is a bill somebody else pays.

The clock on any of this is longer than anybody admits. Getting to a proper depth in a language takes a year or two even when you arrive already expert in another one, and going from PHP to Perl, which share a good deal of ancestry, still took “five years to become really good”. A one-week course buys you an if statement and a for loop. It does not buy the frameworks, the libraries or the quirks, and the distance between those two states is where most of a career actually goes. Noticing that the three of us were discussing a decade of this as if it were a sensible unit of time produced the only correct response available: “my god we’re old”.

Everybody recommends a language they dislike

The list at the top has R in second place. Two of the three people who broadly agreed with that list do not like R, and the third dislikes the syntax and stays for the pictures.

“R is a horrible language. I hate R so much.”

The objections are specific and none of them is snobbery, or not much of it. One of us uses it rarely enough that every return means relearning it, and has settled into a working practice of asking colleagues for their code and modifying that instead, on the honest grounds of “I don’t have headspace for that”. Another can name the mechanism. R gives you at least two ways to say the same thing — the native idiom, and the chained style where a data frame goes through a series of verbs — and lets you mix them freely, and then the libraries disagree among themselves about naming, so one thing saves, the next saves a figure, and the third does something else again. The verdict is Adrian Averin’s*. R’s ultimate problem is “the sum of its small madnesses”. No single one of them is fatal. There are simply too many of them to hold in your head at once.

And then everyone recommends it anyway, because of ggplot and ggtree, because you cannot get those figures out of anything else without drawing them yourself. That is not a recommendation of a language. It is a recommendation of about four libraries, with a language attached that you have to accept in order to reach them.

The case for Python is made on the same terms, though it sounds like a case about the language. Python gets defended for having “there is one way to do it and that’s it”, for a style guide that tells you what to call things, for linting that will hold a beginner to it. The complaint against Python, when it comes, is that it has been drifting away from exactly that — there was a stretch where you could format a string four separate ways, and “that gets my hackles up” — and the fear underneath the complaint is not ugliness, it is R. The fairest line is the flattest: it is everywhere, and “it’s not particularly good at anything”. Try multi-threading in it and the friendliness stops. But a lab technician with an Opentrons liquid handler sat down, wrote Python scripts and had the robot moving liquid inside a few hours, “a phenomenal thing to do”, and that is the entire argument in one afternoon.

The tax on being first

Asked the inverse — what should somebody early in their career avoid — the answer is the trendy ones, and explicitly not because they are bad. It is because of the job you inherit by arriving early.

“You don’t want to be the one to write the first FASTA parser in this language.”

That sounds like an afternoon, and the gap between how it sounds and what it is may be the most useful thing in either recording. A FASTA parser for the files you generate is genuinely an afternoon. A parser for the files that thousands of people around the world generate is a different object with the same name. Take FASTQ, which everybody will tell you is four lines per record: identifier, sequence, separator, quality string. Write the obvious reader — take four lines, emit a record, repeat — and it will work on everything you own. Then a file arrives with the sequence wrapped across several lines, which is legal, and your record boundaries are quietly wrong from that point down. Then one arrives holding “a single Nanopore read which is a million bases long”, and every assumption that a line is roughly a screenful goes with it. Neither of those throws an exception. Both hand you an answer.

And FASTQ is the easy one; SAM and VCF are harder again. The people saying so have done the work. One of us has “written GenBank parsers in three different languages” and will tell you they each work for the job they were written for and “as a general solution they don’t work”. Another wrote both a BioPHP library and a BioJS library, and describes the experience with barely a trace of nostalgia, most of it acquired since.

There is a second charge on that bill and it is subtler. A language that is still moving changes between releases, so the tutorials rot, and one six months old may simply not run. An experienced programmer reads the error, guesses the tutorial was written against the previous major version, and carries on. A beginner cannot, because the diagnosis needs precisely the knowledge the tutorial was supposed to supply — “you won’t have that in your mind’s eye”. So you are not slowed down. You are stopped, and you cannot tell whether you are stopped because you are wrong or because the internet is out of date, which is a much worse place to be than behind schedule.

The money argues the other way for about a minute. Trendy languages pay well. But the best paid of all, on a report one of us remembers reading, are COBOL and Fortran, which are “old and janky” and which run railways and hospital systems nobody dares move. Not relevant to bioinformatics, it was conceded. Filed anyway.

Where a language keeps its libraries

A fortnight later the three of us recorded a conversation about installing other people’s software, and it contains an answer to a question the language episode never quite asks: where does the runtime put things?

Perl’s answer, and R’s, is that everything goes in one directory. A library lands wherever the library path points, with no version anywhere in the name, and the interpreter finds it there. Install a newer version of the same library and it overwrites the old one, and the software that needed the old one is now broken with no way of saying so. “The different concurrent versions are not available.” A system-wide pip has the identical flaw, which is where virtual environments came from — not as hygiene, but as a workaround for a runtime that can hold exactly one version of anything at a time.

The failure has a shape and a name. A diamond dependency is when your tool needs two things which both need a third, and they want incompatible versions of it, and “there’s actually no reasonable way to resolve this”. What you get is a stack trace pointing at a function that exists in one version and not the other, half a day gone, and in the bad cases a conflict with no resolution at all, where the only route through is to open somebody else’s module and edit it until it stops arguing. The point of view of the person this is happening to is the correct one.

“I just wanted to align some amino acid sequences.”

It is worse in this field than in most, because the chain is long and the links are ours. “We call it academic quality for a reason.”

Rust, of all things, comes out of that comparison well, by doing the ugly thing deliberately: fetch every version anybody in the tree asked for, compile the lot into a bloated project directory, and hand back something that “gives you a very reliable executable at the end”.

But none of that ever appears in an answer. Nobody has been asked what language to learn and replied with a description of how the runtime resolves versions, and yet the difference between those two designs is the difference between an afternoon and a fortnight, repeatedly, for as long as you keep using the thing. A good deal of what a language costs you is not in the language.

Where it stands

This argument has been on the record three times. Once in early 2020. Then again two and a half years later, on the stated grounds that it mattered enough to revisit, that the answer deserved updating and expanding, and that minds might well have changed. Then a third time in June 2023, when the second one was published again, unaltered, down to the last pause.

Nothing had needed changing, which is either a compliment to the advice or a remark about the field.

What did move between the first two is Perl. In 2020 the question was still Perl or Python and there were upload statistics to fight over. By late 2022 nobody is putting Perl in front of a beginner, and two of the three still reach for it the moment they are handed a file and asked to pull some numbers out of it, because it is the language they think fastest in, and thinking speed is not a thing you can recommend to anybody else.

JavaScript moved as well, and it is the clearest case of a language being rehabilitated rather than overtaken. The objection to it was never the syntax, it was that the old code would not sit still: arriving at jQuery from Java, “your scope was leaky”, and variables turned up from a hundred lines away with nothing to say where they had been.

“It’s pure, unadulterated chaos. It’s just Mad Max kind of programming languages.”

Then the standards caught up and the frameworks got stricter, and the same person who four years earlier would have told a beginner to stay well away now allows “yeah okay they can give it a go”.

What has dated is the trendy list itself. In late 2022 the examples of a language too fashionable to start with were Rust, Haskell, Go and Ruby, and the warning was that learning something not commonly used limits what you can do here. The mechanism was right and one of the instances has since expired: the parsers got written, the libraries exist, and tools people depend on now ship as Rust binaries with nobody treating the choice as exotic. The warning was not wrong, it was perishable — and the way you know a language has crossed the line is not that it has become popular. It is that somebody else has already written the FASTQ reader.

The one thing recommended without a single reservation, by all three, is the thing none of them counts as a language. SQL keeps surfacing and keeps getting waved off as not really programming, a different mode of working, and then everybody admits to using it far more than they expected to. What it teaches is not syntax but the shape of data: what a primary key is, what a unique field is, how two tables are related. It is old, stable and near enough identical across flavours, and “I actually find it kind of comforting” is not a sentence said about anything else in either recording. Push it hard — tens of millions of rows, multi-table joins, “indexing of indexes” — and the skill turns out to be real, the difference between a lookup taking milliseconds and taking seconds. Push it gently and it leaks into ordinary life: to cross-check two lists you go through the ten-page book first and then check the thousand-page one, because the other way round means reading a thousand pages. That is a query plan, and it is also just how to look something up.

The episode ends on a riff nobody planned. In Python, everything is a dictionary. In Perl, everything is a hash. In Java, everything is an object. And from the one who has been through about eleven of them, “in computer science, everything is a network in my opinion”.

Four descriptions of the same job, each of them true inside one language and slightly false everywhere else. Whichever you learn first is the one you will keep thinking in, and you are almost certainly not going to be the one who picks it. Somebody in the building already did.

The Formats Nobody Designed

“I knew 454 was dead when I saw one of the sequencers in a car park underground covered in dust and I was like yeah I think it’s gone now.”

Not announced in a press release or wound down over a product cycle. Parked. No reagents, nothing left to run on it, an instrument that had rewritten what a genome cost now serving as a very expensive shelf.

The machines go like that. The formats they produced do not.

Somewhere in a mail archive there is still the email that phrap came in, saying “here’s the package please don’t share it”. Getting the thing installed was, in a phrase we all recognised the moment it was said, “sort of a rite of passage”. The folder paths were hard-coded per project, so a day of your life disappeared before you could align a single read.

“I still have the email though just in case I ever need to install it again.”

That is the joke, and it is also the point. The sequencers are landfill. The file formats are load-bearing.

Nobody is in charge

For most of the things we spend all day reading and writing, there is no standards committee. There is a rough consensus, arrived at by everyone using the same handful of tools, which hardened into fact without anyone signing anything. Formats get “defined by what the software actually outputs or what software reads in” — if a popular program emits it, that is the format, and whether it is written down anywhere is a separate and largely academic question.

Annotation is where you can watch it happen. GenBank flat files ship with a disclaimer saying they are not the thing you are supposed to be parsing, which has stopped nobody, because a GenBank file is what got displayed and so it is what everybody had to work with. And the GFF3 that this field exchanges — annotations, then a line reading ##FASTA, then the sequence — went round the world stapled to the output of one very widely used annotation tool. We didn’t agree about where that layout came from. One of us had it down as “an add-in that Artemis put in for the crack”, a habit of the Artemis genome viewer that everyone else copied; another had always assumed it was part of the format. The specification sides with the second: ##FASTA is a directive in the Sequence Ontology’s GFF3 spec3. Most of us still met it in that one tool’s output rather than in the document.

This gives the field a peculiar property: its most durable artefacts were mostly weekend projects. SAM was pulled together by Heng Li at the Sanger Institute over a couple of weeks, published in 2009, and has barely changed since1. Sixteen years later it is still the substrate under nearly every alignment any of us will ever look at.

FASTQ has an even thinner origin story and carries even more. It is “a glorified text file that has the nucleotides that came off the machine and the confidence the machine has when they called it” — which remains the best definition anyone has managed, and is the whole of it. Bases, then how sure the instrument was about each one. An ad hoc format “with such humble beginnings”, and now the thing you go back to when someone asks how you did your analysis three years ago.

Nobody ever wrote it down. It was made at Sanger, used internally, and picked up by everybody else because it was what was there; the name usually attached to it is Jim Mullikin’s, and it is attached loosely. A description did eventually turn up, in a paper by Peter Cock and colleagues in 20102, and the interesting part of that paper is who wrote it. “These guys had nothing to do with actually creating the format.” They had simply spent years parsing it, so they sat down and set out what it was, observing on the way that there was no official description of it anywhere and that “the closest thing was what was on the BWA website and that wasn’t very complete”. It is the citation everyone uses, for want of anything else to cite: a specification arriving a decade late, written by the readers.

The alternatives were not obviously worse. They were less lucky. 454 produced SFF, which encoded the actual flow intensities off the machine as well as the base calls, and was arguably more faithful to what physically happened. Nobody has seen one in a decade. PacBio shipped HDF5, a perfectly respectable container for enormous datasets, and in practice it was miserable: a pain to get anything in or out, because you were stuck with their bespoke tools and the whole thing was “kind of a black box”. Illumina machines still emit BCL files and essentially nobody looks at them — none of us has ever met a person who does anything with a BCL other than immediately convert it to FASTQ.

Good formats do not win. Convertible ones do, and everything converts to FASTQ.

The quality string, done properly

Here is the part worth slowing down for, because it is where the ad-hoc-ness stops being charming and starts costing you data.

The scores are Phred scores: a log-scaled estimate of the probability that a base call is wrong. Twenty means a one-in-a-hundred chance of error, thirty means one in a thousand. Sensible, well-founded, and the least interesting part of the design.

The interesting part is that somebody decided to store them as ASCII characters rather than numbers. One character per base, so the quality string lines up under the sequence and the file stays readable in a text editor. The reasoning, as far as anyone can reconstruct it, was that this was “probably easier to encode it than just having these long strings of integers”. No committee, no specification. A reasonable decision made once, by someone, and inherited by everybody.

It is genuinely nice to read, once you have stared at enough of them. Bad data looks bad: “it gives you a bunch of pound signs and exclamation point so it looks like it’s cussing at you that it’s so bad”. Good data resolves into orderly letters. After enough years you stop parsing and start seeing — “I’ve been like a character in the matrix”.

Then there is the catch, which cost the field a great deal of quietly wrong analysis.

To turn a character into a number you subtract an offset. Sanger-style FASTQ uses 33, so quality zero is an exclamation mark. Early Illumina pipelines used 64, so quality zero is an at-sign. Same extension, same layout, same everything, and every score wrong by 31 if you guess wrong. Worse, the failure is silent and flattering. A Phred+64 file read as Phred+33 does not crash. It reports implausibly wonderful quality and sails through your filters.

The cussing is a clue, up to a point. Exclamation marks and hash signs sit well below the at-sign where Phred+64 starts, so seeing them means the offset is 33. Not seeing them tells you much less. With old Phred+64 data, “you wouldn’t see your curse letters and you wouldn’t be able to tell you see the capital letters and think everything’s fine”. A file that never cusses at you might be excellent Phred+33 or ordinary Phred+64, and being polite is not the same as being right.

Modern data is all Phred+33 and this has stopped mattering, which is precisely why it is worth writing down. It is the sort of thing that vanishes from collective memory about five years after it stops biting people, and then re-bites whoever next opens an archive from 2010.

A digression about squiggles

Nanopore, being newer, got a chance to name things afresh. It produces raw ionic current traces, and the community settled on calling them squiggles. There is tooling with the word in the name. It is, by any reasonable measure, an accurate description of what the signal looks like.

“I refuse to say it. I call them traces because I can’t say squiggle with a straight face.”

The defection was immediate and unanimous.

This is not really about squiggles. Every generation of the technology arrives with its own vocabulary, and some of it sticks while some is quietly declined by the people who have to say it out loud in meetings. Sanger gave us traces and chromatograms, which sound like instruments. Nanopore gave us squiggles, which sounds like something a toddler did. Both describe a wiggly line standing in for a physical measurement. Only one of them lets you sound like you are describing a serious result, and the profession votes on these things without ever holding a vote.

For what it is worth, FAST5 — Nanopore’s raw format — turned out to be HDF5 underneath, the same container PacBio used. It “should not be confused with fast and the furious five”. Naming was never going to be the strong suit.

CRAM exists and is genuinely better

CRAM exists and is genuinely better. When we recorded this in 2019, lossless CRAM against a reference was saving forty or fifty per cent over BAM, and it can go further if you’re willing to throw some data away. At population scale that is the difference between an affordable archive and an unaffordable one. Everyone agrees it is superior. Almost nobody sends you one.

The problem isn’t the file. “It always comes back to the problem that you have to then give it to someone else and then they have to figure out how to work with this file and so you always wind up coming back to FASTQ because that’s just what everyone” — knows. The format is not chosen by the person who understands compression. It is chosen by whoever has to open it at the other end, and that person has been away, and their pipeline was written in 2016.

Which is a shame, because CRAM is the one format here that somebody set out to make. It also owes something to a competition: Sequence Squeeze, in 2011, an international scramble to compress short-read data with entries from Russia, the United States and the United Kingdom. James Bonfield won it from a windowless basement at Sanger, and that work fed into CRAM. “They literally stuck all the developers in the basement at the end of the car park.” The first working implementations were built side by side, in Java at the European Bioinformatics Institute and in C at Sanger, two buildings about twenty metres apart, half in collaboration and half in competition. The C one won on usage. But “having two different implementations of one specification” meant the little issues were teased out quickly, by people who had read the same document and built something different.

These formats, at least, do have a committee. When we recorded these episodes in 2019 and 2020, the Global Alliance for Genomics and Health was already stewarding SAM, BAM, CRAM and VCF, with an actual specification, and without it, we reckoned, every PhD project and every instrument vendor would have come up with its own alignment format. Nanopore has moved on from FAST5 to POD5. The direction of travel is towards things being written down.

Somewhere there is still an email with phrap attached, and somewhere there is a 454 machine under a car park with dust on it. Only one of those is coming back if it has to.

Notes

  1. Li H, Handsaker B, Wysoker A, et al. (2009) The Sequence Alignment/Map format and SAMtools. Bioinformatics 25(16):2078-9. doi:10.1093/bioinformatics/btp352
  2. Cock PJA, Fields CJ, Goto N, Heuer ML, Rice PM. (2010) The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. Nucleic Acids Research 38(6):1767-71. doi:10.1093/nar/gkp1137
  3. The Sequence Ontology, Generic Feature Format Version 3 (GFF3), which specifies ##FASTA as a directive marking an embedded FASTA section at the end of the file. https://github.com/The-Sequence-Ontology/Specifications/blob/master/gff3.md

How This Field Got Here

The 15th of September 1989, a medical school lab in London, and Graham Thorne* is typing bases into an Apple Mac. His colleague holds an autoradiograph up to the light and reads them out loud.

What was at stake was whether a set of clones really contained the urease genes of Helicobacter pylori. The antibody work had produced one faint glow round one plaque and no urease activity from the clone that came off it, which could mean anything. So the sequences went into a program called FASTA — a program first, a file format much later — and were searched against a protein library to see whether a bacterial urease looked at all like the urease from a jack bean. It did. Around fifty per cent amino acid identity, on one strand, and flickering between reading frames: a block of sequence lighting up, then nothing, then a different frame lighting up instead.

Which meant the answer was right and the sequence was wrong. So they went back to the film and counted again. Six A’s here — no, five. Four G’s — five. Homopolymer miscounts, corrected by two people arguing over a piece of X-ray film, decades before pyrosequencing made that error class famous, and everything came into frame except a handful of residues at the end.

“Again this is just laughable looking back at what we have to do.”

It got written up, rejected by Nature for not being exciting enough, and published in Nucleic Acids Research1. That is roughly the shape of everything that follows.

Nobody had a word for it

Graham had come to the lab as a clinical lecturer, a medical microbiologist who talked his colleagues into buying an Apple Mac and started trying sequence software on it. The verdict from the molecular biologists next to him, before the urease result, was that “he’s all right but he’s a bit odd, isn’t he, spent all that time mucking about with the computer, what use is that”. Sequence analysis was the thing you did in a corner “when you couldn’t be bothered to do anything useful”. It was not a career path. It was a personality flaw with a desk.

There was very little to train anyone in. Hardly any genes were known when Graham was a student, and his school biology textbook had “no mention of DNA at all” in the index, though that, he pointed out, didn’t mean he wasn’t taught about DNA at school. The university entrance exam he sat in 1977 — the same year Fred Sanger described the sequencing method that took over the world2 — handed him a set of peptide sequences and asked him to assemble them into a longer protein, by hand, under exam conditions. He lined them up and looked for overlaps, which is the technique, and describes doing it as “just like doing a jigsaw”.

The aptitude had come from somewhere else. Latin and French at school, Esperanto taught to himself, enough hours spent on strings of letters that a page of peptides did not look strange — parsing a Latin unseen being “a similar kind of process of working out what the parts are and how they fit together”. Which is less of a coincidence than it looks, our fundamental algorithms being “essentially derived from text searching and linguistics”. It ran the other way too. Years later he talked a colleague into treating the editions of On the Origin of Species as a series of genomes and aligning them for the insertions, deletions and edits, and what came out was an online variorum.

Graham didn’t hear the word until two decades after he started doing the work. Malcolm Drennan* used it in the mid-nineties, during Graham’s PhD with him, in the sense of this is a growth area, you should get into it. Before that: “I’d been calling it sequence analysis up till then.”

The other direction is just as strange. One of us came to genomics in 2011 straight out of computer science, a child of next-generation sequencing — “I’m an NGS boy” — for whom pulsed-field gels and phage typing are as foreign as that Apple Mac. Both generations arrived with a gap where the manual should be. The gap is just in a different place.

Fifteen bands and a friend

The problem underneath it all never changed: you have isolates, and you need to say whether two of them are the same thing.

The pre-genomic answers were serotyping, phage typing, plasmid typing, multi-locus variable-number tandem repeat analysis (MLVA), seven-gene multi-locus sequence typing (MLST) and pulsed-field gel electrophoresis (PFGE), and none of them are historical curiosities — several are still the reporting standard. “These are classic techniques which have stood the test of time”, and the reason is worth being precise about.

Take PFGE, since it carried thirty years of surveillance. You digest the whole chromosome with a restriction enzyme, run the fragments out on a gel, and count bands. The enzymes are chosen so that you get roughly fifteen. Two isolates differ by however many of those fifteen bands differ, that number is a distance, a matrix of those distances makes a tree. It is “a pan-genomic measure of the genome”, the whole chromosome contributing, and it had a property genomics took another decade to match: “you can show that to your friend and that can make sense”. A picture that travels, under a protocol everyone runs the same way.

Here is the part that cost the field real money: the transformation does not invert. You cannot take a finished genome and predict its banding pattern. Methylation can block the restriction enzyme, so a site that should cut does not. A phage insertion adds content and moves a band, which can make a strain look more closely related to something it has nothing to do with. Run both on the same collection and one PFGE type can blow apart into three whole-genome clades, or several distinct profiles can collapse into one.

So the switchover was not a migration but a break, with an archive on the far side that has to be reasoned about rather than converted. And in the years when sequencing was still rationed there was a chicken-and-egg problem on top: “you have to know what the diversity is before you run WGS”, whole-genome sequencing being the thing that was supposed to answer it. For the Haiti cholera genomes in the early 2010s the selection was made by PFGE and MLVA type, deliberately spread across the diversity the old methods could see. The superseded method picked the samples for the one replacing it.

Meanwhile UK hospitals were tracking outbreaks by antibiogram — grow it, see what kills it, compare the profiles — which takes us “back to Louis Pasteur days”, and is defended on grounds hard to argue with. An antibiogram reports each drug as susceptible, intermediate or resistant, and two of those three are an answer.

“It works, it’s very definitive apart from intermediate.”

The gel on the office wall

Graham remembers a tea-room conversation from around the time PCR arrived. What if you could reproduce those restriction banding patterns with PCR — put restriction sites on the ends of the primers and amplify the same fragments? Better: skip the restriction sites. Make the primers completely arbitrary, drop the annealing temperature far enough, and something will amplify. They tried it, and it worked. Different E. coli strains, different banding patterns. A new typing method, invented over tea, written up and sent to Nucleic Acids Research, who replied that it was “a bit niche. It’s only for bacteria. We’re not really interested.”

Then the reruns stopped matching. The colleague doing the bench work got a few extra bands the second time round and didn’t think they should press on, and Graham, not that skilled in the lab by his own account, deferred to him. It was put aside. About eighteen months later the same technique appeared in the same journal, applied to plants and bacteria as well, and is now called RAPD3.

The retrospective diagnosis is what stings. Amplicons from one round had probably hung around and contaminated the next, and under clean conditions the method would most likely have given the same bands every time. A photograph of one of those gels is still up “on the wall in my office to remind myself not to let opportunities pass you by”.

The methods were not handed down. They were invented by people with a spare afternoon, and the difference between a method and an anecdote was usually whether anybody stayed long enough to work out why it misbehaved.

When the genomes arrived

Graham remembers the lab crowding round a colleague back from a conference in America: Craig Venter and Ham Smith had sequenced a bacterial genome4. Then a second one, “just to show they weren’t a one trick pony”.

What made it a shock was not the sequencing but the shortcut. The proper way, as practised at the Sanger Institute, was top-down: build a large library, subculture down from it, map the whole genome before sequencing any of it. They skipped the map, sequenced everything in small fragments and let the assembly sort it out.

The Sanger adopted shotgun sequencing quickly. The institutional habit took longer: “we don’t start one stage until the previous stage is completed”. Finish the shotgun, then finish the finishing, then annotate — and annotation meant a person looking at every protein-coding gene, reading its homology results and making the call by hand, for months. Nothing like “just running it through a pipeline in five minutes”.

Which collided with the Bermuda Accords, under which publicly funded sequencing had to be released immediately. So the Campylobacter jejuni reads sat in public, unanalysed, while a research community that wanted them now was told to wait a year or two for the annotation.

Graham went round it. Blast the released reads against the E. coli proteins, load the hits into a relational database, wrap a web front end round the lot so people could look up glycolysis or the Krebs cycle themselves. The programmer he hired to build it was an eighteen-year-old on a gap year before medical school, taken on for the best part of a year. “I thought it might take him three months or six months.” It took him about three days. He was called Rich Haddon*, which meant nothing to anybody in the building at the time.

“I can’t compete with youth.”

Seek simplicity, and mistrust it

The noughties opened with dozens of genomes and no idea what was in them. Even for E. coli, studied for the best part of a century, half its genes were unknown until somebody sequenced it.

For Graham, the tool that opened that landscape was PSI-BLAST5, which searches, builds a model out of what it finds, and searches again with the model, so each round pulls in things the last one could not see. What it became for him was “the equivalent of crack cocaine”, and it is not obvious he was joking.

The request had been modest. Jon Ashby*, in Dublin, asked Graham to find all the substrates of the enzyme sortase in the Staphylococcus aureus genome, where they sit scattered around the chromosome carrying a recognisable motif. Graham found half a dozen or more new ones, Jon patented them, and they wrote it up. Then Graham ran the same search across every other genome available, which nobody had asked for, and the tidy picture came apart. Other organisms had several sortases. Their substrates sat clustered with the enzyme rather than scattered. In some actinobacteria the system was not just pinning proteins to the surface, it was building fimbriae out of them.

Graham had a similar experience with ESAT-6, which arrived as the antigen behind a new T-cell test for TB. Asked what the bacterium actually makes it for, the immunologist who brought it in was straight about it.

“I don’t know I’m an immunologist I don’t care.”

The homology searches put the family in Staphylococcus aureus and in streptococci, and conspicuously nowhere in the gram-negatives at all. Twenty years on, what ESAT-6 itself does was “still not entirely clear”, as Graham put it. It seems to have several jobs, probably including one in its own secretion system.

The line that gets quoted about this, from Alfred North Whitehead, is that the goal of every natural scientist is “to seek simplicity but mistrust it”.

Which lands hardest on the organisms everyone had been staring at. E. coli K-12 was, in Graham’s phrase, “handed to us by god” as the model organism. It is nothing of the sort. It carries a type III secretion cluster that frame shifts and deletions have rendered non-functional, and a pair of flagellar genes with no start codons and no promoters — “it’s just like this weird little scar”, all that remains of a second flagellar cluster. Most E. coli carry the same scar, and only a few still have the cluster intact. That is evolution leaving vestigial genes, the way Darwin described vestigial organs, and not the laboratory’s doing.

The laboratory did its own damage. The type strains had been sitting in laboratories for decades being subcultured, and K-12 had been irradiated somewhere in its ancestry. The fix, pushed by Malcolm when he moved to a major sequencing centre, was to sequence fresh isolates alongside them, taken from the wild recently and subcultured as little as possible. It was, in Graham’s analogy, “a bit like the first human genome was largely Craig Venter’s genome”, and “we don’t want Craig Venter’s genome to be the only representative of the whole of humanity”.

Fresh material was not free. On one project the organism would not grow in a lab at all, and eighteen months went on coaxing enough biomass out of it to sequence — longer than the sequencing.

Where it stands

By the end of the 2010s the switch had been thrown. PulseNet had grown past eighty US laboratories, Listeria moved to whole-genome sequencing in 2013, when around 1,600 isolates a year were expected, and Salmonella, running to tens of thousands, followed.

What did not get solved is everything around the sequence. UK samples cluster on Colindale in north London — “a lot of UK samples all appear to come from Colindale in London” — and it is “obviously rife with disease” only in the sense that it is a public health headquarters being used as a default location. Food tested and rejected at a port gets filed under the country that rejected it rather than the one that sent it. The submission form has a field that says source, and nobody can tell you which source it means. Errors found during analysis get fixed in the paper and not in the record — “oh but it’s in the paper you just read the paper” — which does not survive contact with a couple of hundred thousand Salmonella records: “I’m not going to go and read the paper for each one of them.” And for the old strains there is either “a magic technician or a magic postdoc somewhere who has all of this in the back of their head”, or a fifty-year-old paper saying the strain was sent by someone, after which “it magically appears in the world and that’s the type strain we all use”.

Not all of the thinness is carelessness. Listeria is rare enough in the United States that a genome carrying a city and a month could be lined up against a news story from the same city and month, and “we would have linked that patient back to that genome”. So the first release of those isolates gave the country, an anonymised identifier and almost nothing else, with serotype, year, region and an age band in ten-year bins following six months later. Colindale is the same instinct applied bluntly, a default location standing in for a real one.

The prediction on the table in November 2019 was that one central database serving the whole planet would stop being tractable — that everyone would run a fixed, versioned, containerised pipeline locally and push the results up to “some magic cloud place in the sky” to ask whether anybody else was seeing the same thing.

Graham had a mid-nineties prediction of his own to look back on. In a medical journal, he’d argued that email beat post because “there’s no such thing as junk email” — with the correction supplied years later in his own words, that “it’s been very much a double-edged sword since then”.

The workshops have not changed. When Graham and Rich ran their internet workshops for medics, people turned up who “didn’t know how to use a mouse”. Everyone who has run a training course since recognises that room.

It is the one thing in fifty years that nobody has had to update.

Notes

  1. Clayton CL, Pallen MJ, Kleanthous H, Wren BW, Tabaqchali S. (1990) Nucleotide sequence of two genes from Helicobacter pylori encoding for urease subunits. Nucleic Acids Research 18(2):362. doi:10.1093/nar/18.2.362
  2. Sanger F, Nicklen S, Coulson AR. (1977) DNA sequencing with chain-terminating inhibitors. PNAS 74(12):5463-7. doi:10.1073/pnas.74.12.5463
  3. Williams JGK, Kubelik AR, Livak KJ, Rafalski JA, Tingey SV. (1990) DNA polymorphisms amplified by arbitrary primers are useful as genetic markers. Nucleic Acids Research 18(22):6531-5. doi:10.1093/nar/18.22.6531
  4. Fleischmann RD, Adams MD, White O, et al. (1995) Whole-genome random sequencing and assembly of Haemophilus influenzae Rd. Science 269(5223):496-512. doi:10.1126/science.7542800
  5. Altschul SF, Madden TL, Schaffer AA, et al. (1997) Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Research 25(17):3389-402. doi:10.1093/nar/25.17.3389

Early Days of MLST

Before multi-locus sequence typing — MLST from here on — if you wanted to know which lineage of meningococcus you were holding, you put it in the post to Oslo.

The technique it replaced was multi-locus enzyme electrophoresis, which meant mashing up cells with the enzymes still working, running the extract through a starch gel, and staining it with compounds that changed colour to show how far each protein had got. What you scored was mobility. It was the first time anybody could assay molecular variation across a whole population, it goes back to the late sixties, and in Ken Ahlborn’s* summary — he was doing it, at the bench, for a PhD — “it was an absolute nightmare to do”.

It was also not portable. An allele could only be scored by running it on the same gel, side by side, against strains whose alleles you already knew, which makes the gel the instrument and the person running the gel the calibration. For Neisseria “there was one person in the world that could really do” the scheme — Sylvie Marchand*, in Oslo — so the strains went to her. Fifteen loci, one starch gel, one pair of hands, everybody else’s meningococci in the mail.

Seven, because seven was what you could afford

The 1998 Neisseria paper had eleven loci in it, not seven1. Six were the boring metabolic genes anybody would now recognise as classic MLST; the other five were outer membrane proteins and similar, chosen because they varied a lot. Half of one thing and half of another, and the five variable ones “were thrown out because actually they weren’t bringing anything to the party”.

Which tells you what the design brief actually was. It was not find the genes that best describe this population. It was recover the lineages the starch gel has already named, because those lineages were established, people talked in them, and a new method that disagreed with them would have been a new method nobody used. Six genes almost got there. A seventh went in specifically to separate one group that the six could not tell apart, and then it stopped, because the scheme was calibrated to fit the gel and to cost no more than the gel had cost. Seven is not a number about bacteria. Seven is “the trade-off of expense versus resolution” with a Sanger sequencer on the other side of it, and five hundred bases per locus because five hundred bases is how much clean sequence you got reading the fragment from both ends.

It is easy to underestimate what the expense was. Two of the seven Neisseria loci, adk and recA, were Ken’s entire PhD. There were no genomes to read a gene off, so “you had to clone the gene to get the sequence and that took a year”. A year, for one locus. Seven loci is not a modest scheme; seven loci is several people’s PhDs.

The genes themselves were picked on the same logic. “They’re picked to be boring” — under purifying selection, central metabolism, nothing interesting happening in them, so that what you were indexing was ancestry rather than whatever the immune system had been doing to the surface of the cell lately. This worked well enough that it is still how everyone thinks about typing loci, and it was done blind: the genes were chosen before there were genome sequences to choose them from. One of the pneumococcal loci may be present in two copies in some genomes. Nobody could have known.

One base or a hundred, it does not matter

Here is the part that surprises people who arrive from genomics, and it is the part worth getting right.

You sequence your seven fragments. For each one you look up the sequence in a database and get back an integer. Identical sequence, same integer. One base different, different integer. A hundred bases different, also just a different integer — no closer and no further than the one-base case. The seven integers together make a profile, “like a telephone number” in Ken’s phrase, and the telephone number itself gets a number, which is the sequence type (ST).

The flatness is deliberate, and it is not laziness about the analysis. Bacteria recombine. When a locus changes by recombination it does not change one base at a time; it takes in a patch, and the patch might carry one difference or twenty. So the number of differences between two alleles needn’t tell you anything about the number of evolutionary events that separate them, and a distance built out of base counts would be a distance built out of patch sizes. Better to refuse the question. Two alleles are the same or they are different, and nothing in between gets recorded.

The cost of that decision shows up the first time somebody sends you a profile with a gap in it. You need all seven. Six alleles and a blank is not a partial answer, it is not a sequence type, and there is no defensible thing to do with it. People still turn up asking, and the answer is not a kind one.

“You can’t do anything. It just didn’t work.”

The other thing the digital format bought was a database that anybody could reach. A gel is local knowledge; an integer is not. This was the first time epidemiological data had been put somewhere that a hospital in Australia could sequence seven fragments, look them up, and find out that afternoon where its isolate sat in the global picture. In the late nineties that was a startling use for the internet, and it is the reason the sequence-type nomenclature spread as far as it did before anybody had a genome.

The magnifying glass

The standard way to draw MLST data at the start was a dendrogram, and by about a hundred strains the labels stopped being readable. The response to this, in the office of Ken’s supervisor, was a magnifying glass kept in the desk drawer — “he literally had a magnifying glass in his desk so he could read all the labels”.

Declan Hatherell’s* magnifying glass is the origin of eBURST2. If the drawing is unreadable, and the drawing is also wrong — a dendrogram puts the ancestor at the same level as everything descended from it, which is exactly the relationship the data is about — then draw something else. Take really closely related isolates, where you can still see individual events, because “if you look at really, really closely related stuff, everything’s a lot easier”. Find the genotype with the most single-locus variants around it, call it the founder, put it in the middle, and hang the variants off it as spokes.

“There’s nothing clever there at all.”

It got people thinking in circles rather than in dendrograms, which turned out to matter more than the algorithm did. Minimum spanning trees showed up in commercial software not long after, doing broadly the same trick, under a name that has always slightly irritated Ken — there are enormous numbers of minimal ways to connect a set of genotypes, and the thing that picks one of them is not the minimality, it is the founder model. Anyway. The picture won. People still draw those constellations today over core genome data, seven genes having long since stopped being the input.

Whether the clones were ever there

None of it works if bacterial populations do not fall into discrete lineages, and in the mid-nineties that was a live argument. One camp had bacteria as essentially clonal, dividing into neat groups. The other had panmixia, everything so shuffled that lineages do not exist at all. The paper that brought it to a head sorted populations along that axis using the only data there was, which was starch-gel data, this being before MLST existed3. The gonococcus came out at the panmictic end, with “no clonal lineages at all” and everything equally distant from everything else. Meningococcus came out in the middle. What the data eventually said was both: “a soup of recombining things where alleles are just flowing backwards and forwards”, with “croutons of clonal complexes” floating in it, and the croutons were the virulent ones, which is why they were the ones anybody had been looking at.

So a scheme that assumes clonality worked, on the organisms it was pointed at, partly because those organisms were picked for being the interesting ones. In Ken’s experience Staphylococcus aureus is beautifully behaved, and Klebsiella and some of the vibrios are all over the place. The honest summary came from his own thesis, which ended with the observation that there’s no pattern in bacterial population structures — “the closing words of my PhD thesis was something along the lines of there’s no pattern” — and nothing he has seen since has surprised him.

Inês Ferreira* met both ends of that in a single species, Staphylococcus pseudintermedius, during her PhD. Among the resistant isolates she kept seeing the same few clones, ST71 above all. Among the susceptible ones she could rarely find the same sequence type twice; nearly every isolate she typed came back as a new one. Part of the blame, she thought, lay with the scheme, which had started out with only five genes before it moved to seven.

What it turned into

EnteroBase is what happens when you keep the vocabulary and throw away the constraint. It began as two things stuck together: a comparative genomics site with pre-computed analyses, and Mark Achtman’s MLST databases for Escherichia, Salmonella, Yersinia and — for reasons nobody has since been able to reconstruct — Moraxella, a respiratory pathogen sitting in a database named after the gut.

Everything under the hood was rewritten. A cron job pulled new reads off the Sequence Read Archive as they appeared, ran QC, assembly, annotation and every typing scheme in the building, and put the result somewhere searchable. The core genome scheme for Salmonella was built by taking a hand-curated reference annotation, expanding it into a pangenome panel against every decent reference genome available, and then asking which genes were reliably there. Not all of them: the cutoff is 98 per cent, because if you insist on genes present in every assembly in the set, “you basically have three”. Assemblies are bad in uncorrelated ways and a hard threshold just measures the worst one. What came out was 3,002 loci, and the two on the end were noticed.

“Those two are very important.”

Version two arrived a couple of years later, when the database had doubled and some of the original picks had started to break in lineages nobody had sampled the first time round.

The other half of the work was the metadata, and it is the part users actually download. A classifier Zhemin Zhou wrote parses the free text people put in their SRA submissions and sorts it into categories, and after his code had made a first pass it was bootstrapped by a week of going through thirty thousand Salmonella records by hand. “We’re looking up scientific names for random species of North American foxes” — that is a fox, so that is a wild animal — and then getting stuck, because “I do not know the scientific name of a turkey”. Nor did anyone else in the room. And a chicken is a chicken whether it is alive on a farm, in a bag at a supermarket, or crumbed in a box as a kiev, except that it is obviously not.

By early 2021 the database held 267,000 Salmonella genomes, and by our reckoning something like 30 to 40 per cent of them were Enteritidis or Typhimurium, which is worth saying out loud because people keep downloading all of it and treating the download as a sample of the world. It is a sample of what turns up in clinics and food production. “Just because there’s 100,000 data points doesn’t mean it’s random” is a sentence that has to go on manuscript reviews more often than it should.

One rule inherited from the old databases got reversed. Sequence types used to be minted from Sanger traces that people copy-pasted into a web form, or aligned in a Word document by hand, and a new ST was “a badge of honour for a lot of people to be able to say that they’ve identified several novel sequence types in their study”. EnteroBase will not accept them, because “the amount of error is too great”. You can probably see the fossils of that error: of the sequence types minted from traces, around thirty had still never turned up as a genome by early 2021. Some are genuinely rare — Salmonella from the seventies, isolates from rivers — and some are single-locus variants of extremely common types, which is what a typo looks like twenty years later. It never went into print, because we couldn’t pin it down.

What is left

The scheme that was calibrated to a starch gel is mostly gone from the major labs, and it is not gone at all. When we talked to Ken, PubMLST was still receiving Sanger traces, and reference labs without easy access to whole-genome sequencing were still running seven genes. If every proprietary platform vanished tomorrow the method would still be there, because “the primer sequences are in the paper go for it”.

What actually survived is smaller and much harder to replace than the method. It is the names. E. coli ST131, S. aureus ST398, Salmonella ST313 — you say one of those and a room full of people know what you mean, how worried to be, and roughly what it does. Those names came out of seven genes chosen to match a gel, and they have outlived the gel and the sequencing chemistry. As Ken pointed out, genomics has sharpened those lineages far more often than it has overturned them. The lineages are real. The number of loci was a budget.

“It will be the gift that MLST leaves us really.”

The numbers are not even memorable. The reference genome the Salmonella core scheme was seeded from got called ST131 on the day and corrected to ST313 in the show notes afterwards, by the person who built the scheme, who mixes them up as often as everyone else does.

Notes

  1. Maiden MCJ, Bygraves JA, Feil E, et al. (1998) Multilocus sequence typing: a portable approach to the identification of clones within populations of pathogenic microorganisms. PNAS 95(6):3140-5. doi:10.1073/pnas.95.6.3140
  2. Feil EJ, Li BC, Aanensen DM, Hanage WP, Spratt BG. (2004) eBURST: inferring patterns of evolutionary descent among clusters of related bacterial genotypes from multilocus sequence typing data. Journal of Bacteriology 186(5):1518-30. doi:10.1128/JB.186.5.1518-1530.2004
  3. Maynard Smith J, Smith NH, O’Rourke M, Spratt BG. (1993) How clonal are bacteria? PNAS 90(10):4384-8. doi:10.1073/pnas.90.10.4384

Assembly, and Other Acts of Faith

Take a draft assembly, strip out every header line, join what is left into one long string, and the genome is closed. One sequence, chromosome-length, a perfect N50. Nothing downstream will complain, because nothing downstream can.

It was published as a joke — a short recipe on Tristin Lindemann’s* blog for how to close a genome, performed with the wires showing.

“But that has happened in real life.”

Not that exact recipe, presumably. Aggressive enough scaffolding gets you to the same place by a route that looks like work, and it has been done both deliberately and by accident, and it has been done “for medically important genomes”, and nobody is naming names.

The reason the trick lands is the whole of what follows. An assembly is the one thing we produce that cannot be checked against the truth. The chromosome it came from went into a machine, got shredded on the way through, and no copy survives to compare against. What you have at the end is a hypothesis with a file extension, and every metric, polishing round and careful look at it by eye is a proxy standing in for a comparison nobody can make.

Nobody ordered forty pieces

You sequence an isolate, you assemble it, and what comes back is not a chromosome. It is forty pieces, or a hundred, or four hundred, in no useful order, with nothing to say which end of which joins to which. The first time it happens it reads as a failure — as though more coverage or a better assembler would have produced the single sequence you were expecting. It would not have.

The methods have turned over about three times getting even this far. Before the earliest greedy assemblers the state of the art was printing the reads out and overlapping them on a desk by hand; CAP3 automated that, and you still went over the result by eye. Then 454 and overlap-layout-consensus, comparing every read against every other read, which is fine while there are not many reads. Illumina and SOLiD ended that. Coverage went up, all-against-all stopped being affordable, and the field moved to graphs. The move was made concrete by an assembler called Euler, out of Pavel Pevzner’s group1, which went on to write SPAdes.

The costs of the era are hard to convey now. Benchmarking an assembly improvement pipeline meant running AMOS on one core against one bacterial genome — “five million characters in one month”. The same site ran half a million Velvet assemblies through an automated pipeline over the following years, because anything coming off the sequencers that looked bacterial got assembled whether anyone had asked for it or not.

Velvet is the one people remember fondly, and the fondness is not really about quality. It would hand back something reasonable every time and “it works without fail, it just would not crash”, which put it ahead of tools that were cleverer and less willing. It was conservative to a fault, so that “the contigs were almost never wrong” and two genes sitting on one contig really were neighbours, which you could take to the bank. It had its moods. Hand it a library with a sloppy insert distribution and Velvet would just chuck a fit.

By 2020 SKESA was doing a bacterial genome in two to ten minutes where SPAdes took thirty to ninety. That was day-to-day experience rather than a like-for-like benchmark, but either way it was a long way from the month.

Why the graph stops

Here is the plain version. Chop every read into overlapping words of a fixed length — those words are k-mers, which do a great deal of work elsewhere in this book; here their only job is to be nodes. Draw an edge between two words overlapping by all but one base. The structure is a de Bruijn graph. The genome is a walk through that graph, and the assembler’s job is to find it.

Here is the honest version. It is a filing system that deliberately forgets where anything came from. Two identical words, one from either end of the chromosome, are the same word, so the graph holds one node for both. That forgetting is what makes the structure small enough to build at high coverage, and it is exactly what makes the answer ambiguous.

The word length is the one number you have to choose, and nothing underneath it tells you what to pick. The answer offered on air was “read length minus one. As long as it’s odd.” — which is the ceiling rather than the setting. Modern assemblers dodge the choice by running several word lengths and merging the results, and the ones they pick sit a long way below that ceiling.

It is named after a Dutch mathematician, which raises the pronunciation problem. Nobody on the show could say it. Several attempts were made inside twenty seconds, none of them agreed, and one of us tried to triangulate via a football team, which did not help.

“We’re gonna get into so much trouble.”

No roasting followed, but the anxiety is real and general — the field is full of terms that arrive in papers and get read silently for years before anyone has to say them out loud in a meeting. The recommended route out was the reading, and it still is. The Velvet paper is a good place to start2 and Daniel Zerbino’s thesis is better, because it explains “what a bubble is, what a spur is”, and why the word length has to be an odd number — the small structural facts that turn the graph from an abstraction into something you can reason about.

Anyway. Repeats.

A repeated stretch longer than the word length collapses into a single path with several ways in and several ways out. Salmonella carries seven ribosomal RNA operons, around five kilobases each and near-identical. A 150-base read landing inside one of them cannot say which one it was in, so the assembler arrives at that node with seven edges coming in, seven going out, and nothing to pair them up with — five thousand and forty ways of joining them, one of them right. So it stops. That is what the end of a contig is: a place where the evidence ran out, marked by breaking the sequence rather than by guessing.

The consequence is the useful part. Your contig count is mostly a readout of your organism’s repeat content relative to your read length, not of how careful you were. Shigella is essentially E. coli with a plasmid, which is true and does nothing to prepare you for the assemblies. An N50 of thirty or forty thousand is ordinary for Shigella off short reads, against something like a hundred and twenty to a hundred and fifty thousand for E. coli. Same chemistry, same pipeline, same day. The difference is insertion sequences copied all over the chromosome, each longer than a read, each a fork the graph cannot resolve. Fragmentation like that is a property of the bug, and you have to know your bug.

Sequencing errors do the opposite damage. Each one mints a word occurring precisely once, hanging off the graph as a dead-end twig, and at ordinary coverage those get pruned without incident. Push the coverage very high and they stop looking like twigs: at ten thousand fold, “errors are going to start looking like real signal”, the assembler reads them as genuine minority variants, and “the graph will just go crazy”. Which is why subsampling down to something like a hundred fold routinely gives a better assembly than throwing everything you have at it — one of the very few places in this job where using less of your data is straightforwardly the right move.

At every unresolved fork there is then a policy decision, and that is where assemblers actually differ. SPAdes goes for the most nucleotides in a row it can justify. SKESA sees the ambiguity and breaks.

“I’m going to stop right here because I don’t want to make a miscall.”

One in-house comparison came back with zero ambiguous bases across the SKESA assemblies and a non-zero count across the SPAdes ones. Contiguity against correctness, settled at the fork, inherited by everything downstream. That comparison predated SPAdes gaining an isolate mode, tuned for single bacterial genomes, and nobody had rerun it by the time we recorded. SKESA is also deterministic, which most assemblers are not — run one multi-threaded and “sometimes they produce different results”, because the threads seed from different places and finish in different orders. Velvet on a single core reproduced itself exactly; Mira was randomly seeded and gave a different answer every run. It suggests a reviewer’s question almost nobody asks, which is to run the thing again on one thread and show your working. “Might be a third reviewer comment.”

Read healing

Everything before the assembler is an attempt to hand it a simpler graph, and the name one lab gave the whole family of it — trimming, correction, filtering — is read healing, which is a better term than the field deserves.

The simplest trimming is not subtle. The first fifteen bases are usually weak and the last few fall away, so you chop both: “You just do a bit of chopping, there you go.” Most of us trim on quality as well, and there the agreement ends. One of us suspects the choice between a fixed threshold and a sliding-window average matters more than anyone admits; another has used both and can’t see much difference.

Adapter sequence in the middle of a read means a chimera. Quality collapsing partway through a run often means something happened in the building — reagents ran out, or “the machine is paused and people are fiddling around with it”, and the run takes a hit while it restarts.

Read correction is the ambitious one, and the argument for it is graph-shaped and hard to fault: leave the errors in and “your assembly graph is going to be incredibly huge and when you try to traverse it you’re probably going to make mistakes”. Fix them first and the graph stops branching. The bill is that you are deciding, before assembling anything, which of your minority observations are wrong.

“You’re scooping a lot of stuff under the carpet, and you’re sort of cherry picking a little bit.”

The concrete version of that bill came up over lunch on the day we recorded: somebody had lost a three-kilobase plasmid from a Nanopore assembly. The likely culprit was a five-kilobase cut-off, with everything shorter used to correct the longer reads and never assembled in its own right. The plasmid didn’t have to look like an error. It only had to be short.

Every filter here has a version of that. Discard rare k-mers to clean up errors and you can lose the low-copy plasmids; discard over-abundant ones to flatten uneven coverage and you risk the high-copy ones instead. The defence is knowing what ought to be there before you look, which is why Andrew wrote TipToft: it reads the raw data for replicon and incompatibility-type sequences, tells you a plasmid is present, and sends you looking for it in an assembly that may no longer contain it. Hunting for something you know exists is a different activity from hoping.

Checking a thing you cannot check

The reflex answer is N50, and N50 is the metric everybody quotes and nobody can state. One graduate lab spent a day trying to get it into a single sentence and came out with two. Attempting it live went worse. The definition offered first — the length of the contig in the middle when you sort them by size — is clean, memorable, and is the median contig length rather than N50. It was agreed to on the spot. The correct one is length-weighted: sort the contigs longest first, total their lengths, walk down the list until you have accounted for half the assembly, and the length of the contig you are standing on is the N50. That version came last, was accurate, and was abandoned by its own author with “I made it more confusing”.

Which is a fair account of the metric’s career. It rewards contiguity, contiguity is what an aggressive scaffolder manufactures, and the more contiguous half of the world’s assemblies is not the more correct half.

It is worth spelling out what a scaffolder does, because the word hides a lot. Assembly proper stops at contigs. Scaffolding is the pass after that: paired reads say two contigs are physically linked and roughly how far apart, so the pair gets written out as one sequence with a run of N’s standing in for the stretch nobody sequenced, and gap-filling then goes back to the reads to try to turn the N’s into bases. None of that is dishonest. But the metric cannot tell a base from a placeholder, and a scaffolder set loose will hand you one long sequence held together mostly by a willingness to believe a mate pair.

The checks that catch things are biological rather than statistical. Total length first, because a six-megabase Salmonella is a strong claim — “That is impossible. That’s not how this works.” — except that it is not quite impossible, since one carrying big plasmids will run to five and a half or six, and the largest serovar anyone could name off the top of their head was Weltevreden, which one of us had assembled. RefSeq holds a 150-kilobase E. coli and a ten-megabase E. coli, and somebody deposited both.

Then map your reads back onto your own assembly and look for the ones that suddenly point somewhere else, which is the signature of a scaffolding error rather than a biological one. Then classify the contigs, because Salmonella contigs hitting Klebsiella means “either you’ve got problems or you’ve got a Nature paper”. Then, once long reads started producing whole chromosomes, a properly structural check: Socru takes the order and orientation of the ribosomal operons around the chromosome and asks whether that arrangement could exist, against a catalogue covering four hundred and thirty-three species. Seven operons in a Salmonella is expected. One is a collapse, fifteen is a duplication, and an order that cannot be traversed from origin to terminus is a misassembly wearing a good N50. What you want is “a pattern that is biologically legal”, which is a phrase worth stealing for other purposes.

Polishing belongs to the same category of hopeful ritual. The community position on Pilon was to run it four times, at which point “it feels kind of magical”, and the objection from across the table was immediate — “I suppose it can reinforce errors as well”. Repeat an operation until it stops changing anything and you have a fixed point, not a fact.

Where it stands

All of this was recorded in the first half of 2020, with the caveat attached at the time that any of it might be obsolete before anyone heard it, and the standing advice to consult “your physician or bioinformatician if problems persist”.

Some of it moved as expected. N50 was already sliding down the list, because Nanopore had begun handing over whole chromosomes without anyone having to work for them, and that has only become more true. NCBI’s reassembly of the archive with SKESA, which had then got as far as Listeria, kept going, so a great many public assemblies you now download were made by a tool chosen partly because it refuses to guess. The graph is no longer something you have to imagine, either: Bandage draws it, and for a long-read assembly that will not close, the picture generally shows you the repeat you are stuck in. It is what to open “if you’re trying to see why it didn’t circularise”.

By 2024 Nanopore assemblies were being called perfect, and Tristin, on a hackathon panel, wasn’t having it. To him, ten errors left in a genome was still a fail if they were indels rather than substitutions. For a lot of tasks that’s fine, he allowed, but for others “we don’t want 10 frameshifted genes”.

What has not moved is where you stand at the end. You have run the metrics, mapped the reads back, checked the operons and polished until nothing changes, and none of it amounts to a comparison with the sequence you were trying to recover.

“I try not to overthink it, but I know that there’s always something more I can do and probably some mistakes that I’ve made along the way.”

Somewhere in a public database there is a genome whose N50 was decided by a scaffolder rather than by the organism. Nothing in the file says so.

Notes

  1. Pevzner PA, Tang H, Waterman MS. (2001) An Eulerian path approach to DNA fragment assembly. PNAS 98(17):9748-53. doi:10.1073/pnas.171285098
  2. Zerbino DR, Birney E. (2008) Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Genome Research 18(5):821-9. doi:10.1101/gr.074492.107

Sequencing Technologies, and the Arguments About Them

“There’s no point anymore in doing any bioinformatics software for short-read sequencing.”

That was said out loud, into a microphone, to a live audience of scientists at a medical research unit in The Gambia in January 2020, as the opening position in a formal debate on whether Nanopore had rendered onsite Illumina sequencing obsolete. A real motion, with sides: Rich Haddon*, Kerem Yildirim*, Mansur Kamara* and Stephen Hollis*, with one of us making up the numbers, and Penny Ashworth-Clarke* in the chair. One panellist joined over Skype and was lost entirely to bad audio, which for a debate about quality thresholds is at least on theme.

Within weeks of that room emptying, the thing everyone in it would spend the next two years sequencing arrived. The argument got settled in public, at scale, by people who had no time to be having it — and it was not settled by either of the machines on the ticket.

The room did not agree

Stephen’s objection came back inside a minute and it was about physics. You apply a voltage across a pore and read the wobble in the current, and his verdict on where that ends up was flat: “I just can’t see that that will ever become high enough quality.” Not for SNP resolution. Perhaps eighty or ninety per cent of applications could move to long reads, he thought, and there would always be a residue that needed Illumina-grade base calls, or PacBio for anyone with access to one.

He then predicted that Nanopore would probably end up the leader anyway, which is a stranger position than it sounds, given that over ninety per cent of sequencing was Illumina at the time. The reasoning was that fewer and fewer people coming new to sequencing would go down the Illumina route at all, and that accessibility would finish what accuracy could not.

Nobody in the room could produce the study that would have decided it. Had anyone actually tried calling SNPs off Nanopore for a hospital outbreak — a Staphylococcus aureus going through a ward, an Acinetobacter — and shown that it worked, or shown that it didn’t? The question was asked plainly and went unanswered, and the nearest thing to a reply was that the field had lived through this before with the GA2, where the quality was low enough that you threw away SNPs and made judgement calls and got on with it.

From the surveillance side, Samina Qureshi’s* answer was that it depends entirely on what the sequence is for. Her group was running hybrid — MinION and Illumina on the same isolate — and the reason was not fastidiousness. A long read will tell you, went the follow-up question, that a set of salmonellas from Newcastle and London look sufficiently similar to be one outbreak. It will not yet tell you which kebab shop in Newcastle handed it to which person, and “that kebab shop needs to be closed down” is a sentence with a legal department attached to it.

“You might go and put some farmer out of business, and it costs the UK government millions of pounds if they sue us.”

Accuracy, in her line of work, is not a quality metric. It is an indemnity, and it is the part rarely said in a methods section. A triage scheme got floated from the floor — Nanopore everything, sort into baskets by similarity, spend Illumina only on the fifty isolates that cluster — and Samina’s reply was that her group was already doing it, in the other order.

Then the position that has aged best, which came from Mansur, whose building the debate was being held in. Obsolescence is not a property of an instrument. It arrives at different times in different places, and it arrives sooner where the alternative is harder to reach — in his words, “it would be obsolete much quicker in Africa”. Not because the standard is lower — his framing was that “we might not worry about error rates but we can do 90% of the stuff that we want to do in Africa” — but because what is being weighed is not one error rate against another. It is an error rate against a courier, a service contract, a cold chain and a building. From the London side of the same school came the entry price: “It’s $1,000 to get into it rather than having to build an entire building.” Mansur’s own procurement plan was to get the money first and choose the technology after, because “two years is a long time in the sequencing business”.

He had the concrete version of that already. A Gates Foundation grant for metagenomics on sepsis had come tied to a deal with Illumina for twenty small instruments, the money and the platform decided together somewhere else, and it took a meeting in Addis Ababa with Rich in the room to get the terms extended so the work could be done on Nanopore instead. Which is the argument in one procurement document. What ends up on the bench is not chosen by whoever calls the better SNP.

The last word from the public health end of the table was that the first thing wanted from a sample is the pathogen and then the serotype, because that is what decides whether a vaccine gets deployed. At which point a SNP is not in the question at all.

Somebody summarised the state of play as rumours of a death “greatly exaggerated”. The reply was “it’s inevitable”. The recording ends without a winner, which is correct, because there wasn’t one.

A thousand and one bases of spike

Fourteen months later, in March 2021, the fastest route to a variant of concern in Denmark ran through neither machine.

Robin Delmas* and Jeppe Kjærgaard*, at a Danish university, had noticed that the qPCR testing lab was already touching every positive sample1, and that everything needed to read a thousand bases of that sample was sitting in the same freezer. Positive wells get picked into a fresh 96-well plate. An RT-PCR is set up using the same enzyme mix as the diagnostic assay — chosen, and this is the whole design philosophy in one clause, because “we chose not because it’s good, but because it’s accessible”. Take a forward primer from one ARTIC pair and a reverse primer from another and you get 1,001 bases across spike. Then 1.5 microlitres of that product goes to a commercial Sanger provider, unpurified, and the traces come back the following morning.

No purification step. Nothing to pool and nothing to split apart afterwards, and no assembly, because a single read is the answer. It “costs $5 to run a sample from end to end, including all plasticware, including everything”, and produces 300 kilobytes of data. Fastest observed swab to variant call was about thirty hours, average fifty to sixty.

The analysis is the part worth slowing down for, because it is a lesson in matching resolution to question rather than to ambition. The AB1 traces get base-called with Tracy, mapped to the reference with Bowtie2, and then samtools mpileup is run once per position of interest — around a dozen of them, one pileup per mutation. That is not an efficient way to use mpileup and nobody pretends otherwise; it is chosen because when a run behaves strangely you can go and look at exactly one pileup by hand. The whole plate takes about ten seconds, and rather longer in a browser, which we will come back to.

And then the decision that makes the method work, which is a decision to do less. They stopped calling variants. Robin had first tried to make the calls forgiving, so that a lineage could still be called with a mutation missing or an extra one present, but in a 1,001-base window B.1.351 and P.1 barely differ, and a run of American variants of interest looked the same as each other. Jeppe had been through the list on covariants.org and reckoned they could still separate nearly all the interesting ones. Robin’s answer was that this held only if every position mapped and the read quality was high enough, and his summary of trying was “I’ve tried and it wasn’t pretty”.

So the output is not a lineage. It is a table of mutations, one column per position, with a quality value attached to each call, and the naming of things is left to somebody else.

Which sounds like a retreat and is the opposite. The pipeline’s customer is a contact tracer, and what a contact tracer needs is whether to put a normal amount of effort into this sample or an enormous one. An E484K answers that. A lineage name does not answer it any better, and the whole genome is coming off a Nanopore run two days later anyway. The design leans the same way throughout, on Robin’s principle that “false positives are less problematic than false negatives” — an inversion of the instinct every one of us was trained with, and correct here.

It caught Denmark’s first P.1 on a Monday. Not a lineage call — the amplicon carried every hallmark mutation of P.1 and none of the mutations that would have suggested anything else, which was enough. The authorities were told by telephone, by Slack and by any other channel available; the person was isolated; whole-genome sequencing confirmed it two days later. The screen has zero personal information in it and never sees anything but a barcode.

Above about cycle threshold 30 on their setup, which Jeppe reckoned was roughly 35 on other systems, the success rate falls away, and the response to that is the five-dollar shrug: a failure costs five dollars, so try every positive and report only what comes back. Somebody wrote to Jeppe to say “welcome to 1995”, and he repeated it back to us with some pride. The honest version came with it — he hadn’t touched Sanger sequencing since starting his master’s thesis in 2008, and “I was actually a little bit embarrassed at first”.

The one button

The tangent this earns is about where the software ended up, because the same constraint applied one layer up. The screen was going out to hospitals and regional test centres, and a lot of them, Robin pointed out, had nobody on staff to do bioinformatics; Jeppe added that a command line would be a wall on its own. Robin’s usual answer would have been a web service, and a web service means hospitals uploading patient-derived sequence to a website, which is a conversation nobody wanted.

What happened instead was that a small company, BioLib, got the whole pipeline running in WebAssembly — Bowtie2, samtools, a Tracy build, a Python interpreter — so the page you open is the tool, and nothing you load into it is transmitted anywhere. You put in a zip file, you press one button, you get a table. It costs three or four minutes instead of ten seconds, which against a two-day turnaround is nothing, and it removes the step Robin reckoned is most of the job: “dependency handling is the thing that you spend 80% of your time on”, and then you spend the other eighty per cent doing the work.

“You guys are really scaring me with the one button thing.”

That reaction was ours, and it was not entirely a joke. Robin’s answer was that it wasn’t for every pipeline: this one had a single job, and the fewer things there are to twiddle, the fewer ways somebody has to get the wrong result. The rest of the machinery pointed the same way, right down to the paperwork. bioRxiv “said no thank you” to the preprint, so it went to medRxiv, which requires ethics approval — for a study which under Danish law needs none, because it holds no personal data. Which meant asking a Danish ethics board for a statement confirming that a Danish ethics board was not required.

Where it stands

The first crack in the January 2020 prediction shows up inside our own catalogue, at the end of 2022, in a conversation about what people had been demonstrating at that summer’s conferences. Nanopore’s Q20 chemistry was arriving, and the assessment was that there were papers “showing that it works it’s not just hype”.

That is the prediction that broke, and it broke completely. The claim was that a voltage across a pore could never reach SNP resolution. The chemistry has turned over twice since, per-read accuracy went past the point where the objection was aimed, and Nanopore-only bacterial assemblies are now used routinely for exactly the ward-outbreak comparisons that the room could not name a single study of. The question asked from the floor that afternoon — has anyone actually done the comparison — has been answered in print many times since, and the answer is yes.

The conclusion drawn from that wrong premise, though, held up. Illumina is not obsolete. Nor did its share flip, which was the other half of the same person’s forecast — short-read software did not stop being written or funded, and a great deal of the world’s sequence output still arrives 150 bases at a time. So the person most wrong about where the chemistry would end up is also the one who called the outcome, by a route that has not survived, and whoever called the direction correctly was wrong about the year. That is the usual distribution in these arguments.

What held up without qualification was the geography. The most accurate thing said in that room was that obsolete means one thing in Norwich and another in The Gambia, and the entry price is still what decides who sequences their own samples and who ships them somewhere. The one change since then that speaks to any of it is not a chemistry improvement. Among the things worth naming on Illumina’s new flagship at the end of 2022 were reagents that ship at fridge temperature instead of frozen, which is worth nothing at all on a bench in Norwich and quite a lot at the far end of a supply chain. None of that was ever going to be settled by whoever had the lowest error rate.

The rest has dated in a friendlier way. Element Biosciences turned up in a 2022 round-up as a rumour, one instrument that “seemed to be solving all of my problems, which make me very very very suspicious”, and it shipped, and the suspicion was unwarranted. Being spoiled for choice was the verdict at the time, and that part has held. The all-in-one sequencer with the compute bolted underneath got the reception it deserved, which was a story about the original PromethION tower being “out of date and underpowered within about five minutes”, and the current arrangement in a good many sequencing rooms is still a gaming rig — “the most pimped out, over spec computer gaming machines ever”, bought for the GPU, running base calling instead of anything a fourteen-year-old would want.

The number that decides whether a run works, meanwhile, is on nobody’s spec sheet. Ninety-six E. coli went onto Nanopore, the first attempt averaged about one and a half kilobases and did not work very well, the prep was adjusted to go a little bigger, and somewhere around five kilobases “your life just magically sorts itself out”. That threshold is not in a paper. You find it by having a bad week.

Nobody wrote the paper. Somebody ought to write a short one defining the terms, was the suggestion, once it had become clear that three hundred base pairs was being called short when “short was 35 base pairs” — and that the tools would not survive the drift anyway, since “half the software we use for Illumina would just crash and burn” if read lengths moved much further, because somebody hard-coded a limit years ago on the reasonable assumption that nobody would ever exceed it. Short, medium, long and ultra-long still have no agreed boundaries. Read length is now a number you look up per platform per kit, which is what the terminology was supposed to save us from.

The motion was never carried. The five-dollar tube, which nobody in that room proposed, went out to a commercial provider one afternoon and came back the next morning, carrying the answer the two machines were still arguing about.

Notes

  1. Jorgensen TS, Pedersen MS, Blin K, et al. (2023) SpikeSeq: a rapid, cost efficient and simple method to identify SARS-CoV-2 variants of concern by Sanger sequencing part of the spike protein gene. Journal of Virological Methods 312:114648. doi:10.1016/j.jviromet.2022.114648

Trees, and How to Argue About Them

A supervisor at an Australian university, Fraser Dunbar*, had a Perl script that filled in a large configuration sheet, which fed a program called Blast Atlas, which drew a circular comparison figure. He wanted the middle two steps to stop existing. So he put it to his group as an open tender — somebody write me something that takes the BLAST results and just draws the picture — and “the reward was one jug of beer to be the first one to write it”. A jug in Australia is two pints, which was judged light for the work involved.

The work was supposed to take a week or two, and the first version did. Then came the questions. Can it take GenBank as well as FASTA? Can it draw the annotations out of the GenBank file? Can it show coverage from a BAM file? Nine months later there was BRIG, the BLAST Ring Image Generator, published in August 20111, and Nabil-Fareed has spent the years since finding it cited as a method. Figure two, this BRIG analysis shows.

“It’s just Blast. I mean, you’re giving me too much credit.”

Everything in this chapter has that shape. A tool takes whatever you give it, cannot refuse it, and hands back a picture that goes in the paper and gets argued about. None of the arguments is ever about whether the picture came out.

The tree always comes out

Ronan Keogh* always tells people the same thing about phylogenetics software: give it an alignment and it will give you a tree. Not an error, not a refusal, not a note saying there was nothing here to work with.

“You will always get a tree, every program will create a tree, it’s not like you’ll get an error and it says nothing can come out at the end.”

The nearest thing to a complaint the software has is a star — every tip hanging off one node, no structure at all — and, as Ronan points out, “you get a star but that’s still a tree, technically it’s acyclic”. Rafael de Souza Lima* takes it further. Align entirely random sequences that share no ancestry and there’s still signal in the alignment for the tree inference to pick up. The branch lengths will probably be enormous. They’ll also be finite, because the software assumes branch lengths are finite.

So the question Ronan puts to anyone arriving with a molecular epidemiology problem is not which program to use — “the first question always is why are you building a tree”. A fair number of the honest answers are that the paper needs one. If all you want to know is which isolates are closer than which, a SNP matrix and a few cut-offs will tell you.

Ronan thinks the reasons for parsimony and distance methods have mostly gone away; they were there because twenty or thirty taxa used to mean leaving a machine running for days. Rafael isn’t so sure. In an outbreak or a clonal complex you can assume there are no back substitutions, and Ronan concedes the point: over a short time span in a small population, unobserved multiple substitutions are so unlikely that a parsimony or distance tree will probably give you the same answer as maximum likelihood.

Sometimes the cheap answer is the whole answer. In Ronan’s account, it’s doing something with the tree — clustering on it, putting dates on it, working out when transmission happened — that makes you pay for the expensive machinery.

What the branch lengths are made of

The routine bacterial version is a SNP alignment: map everything to a reference, call the variants, concatenate the variable sites, hand that to RAxML or IQ-TREE. It works. It also does something invisible to the numbers that come out: run the same data both ways and “the topology is more or less correct, but my branch lengths are all rubbish”.

The mechanism is worth having straight, because nothing warns you, and it’s the thing Ronan calls his pet peeve. A substitution model has two moving parts: the rates of change between the nucleotides, and the frequencies of the nucleotides themselves. The frequencies are estimated from the alignment you supplied, and a SNP alignment holds only the positions that varied, which makes it a poor sample of a genome’s composition. Ronan works on tuberculosis, where the genome is GC-rich and the variable sites can come out AT-rich, so the model gets fitted to a base composition the genome doesn’t have. The topology mostly survives, because the order of the splits is carried by the pattern of the SNPs and the pattern is real. The distances do not, because they are scaled by a parameter you falsified on the way in.

Ronan runs through the corrections, named for whoever worked them out. The Lewis one recalculates from the variable sites alone. The Felsenstein one takes the total number of constant sites. The Stamatakis one takes how many constant A’s, C’s, G’s and T’s there were, which gets closest, and “almost definitely will give you the correct information, almost”. His advice, after talking it over with Rafael, is to skip all three and put the whole alignment in, constant sites and all.

Except that, as Ronan says, whole-genome sequencing isn’t always whole genomes. It’s whole genomes minus the repetitive regions, discarded by your pipeline long before anybody was thinking about models, so the constant-site count you type in is itself an estimate of something nobody observed.

Which is why the checks Ronan and Rafael run on a tree somebody brings them, excited, are mostly not checks on the tree. Branch lengths first. For Ronan, a species-level tree with tiny branches means something has gone wrong, perhaps a lot of the data cut out by mistake. Rafael looks for branches sitting at ten to the minus six rather than zero, because the software would sooner hand you an epsilon than admit it has nothing to report; that might not be wrong, but it sends him back to the data. Then Ronan’s check from outside the figure entirely, “where the bio of bioinformatician comes in” — if you work on tuberculosis you know which lineages belong beside which, and when they are not, something upstream has gone wrong that staring at the picture will never identify.

An argument nobody wins

The argument people genuinely enjoy having is about models, and the reasons offered for a preference are frequently genealogical. Rafael’s favourite is HKY, which is Hasegawa, Kishino and Yano2, and the first case he made for it was that he did his PhD with Kishino. The reply came back without a pause, GTR being the general time-reversible model rather than anybody’s supervisor: “Yeah, I did my PhD with the general from GTR.”

Rafael had a second reason, offered straight, and it turns out to be the better joke. HKY is the most complicated model that still has an analytic solution, so the machine never has to sit down and solve the eigensystem to get the rates out. The saving, by his own rough estimate, is somewhere between one per cent and two-tenths of one per cent. Ronan’s case for GTR was biographical too: his PhD was on HIV, which needs it, so that’s where he started.

Both are wrong, as Rafael pointed out, which is not a concession but the entire point of a model. It is not a description of what happened. It is a device for not being astonished by two changes at the same site. The best version Ronan has heard is that a model is like the London Tube map, which “should be accurate enough to get you all the information that you need” and no more detailed than that, or it stops working for anything but the one journey it was drawn for.

The naming is worse than the arguing. As Rafael explains, RAxML’s CAT model and PhyloBayes’ CAT model are unrelated ideas that both abbreviate category, and neither of them means concatenate, which was our first guess.

But the base frequencies live inside the model too, and the base frequencies are exactly what a SNP alignment gets wrong. The argument everybody enjoys is about which model; the thing that wrecks the tree is what was fed to it. Rafael’s answer is to run model selection and take what it tells you, though in practice, he says, the choice rarely makes much difference. Ronan sees a difference in temperament between the two programs everyone uses. IQ-TREE tests the models and picks one for you. RAxML, as he described it in 2020, used GTR unless you told it otherwise, and as far as he knew its author had said more than once that you didn’t need anything else.

Their own choices had moved in opposite directions. Rafael went back to IQ-TREE after a dataset took longer on RAxML than he was used to with IQ-TREE. Ronan had used RAxML for years, tried IQ-TREE on a large tree, was told to expect about two months, went back to RAxML NG — and by the end of the conversation was contemplating trying IQ-TREE again.

The R word

The next problem arrives pre-labelled: “I said the R word, recombination, which is absolute poison when it comes to phylogenetics”.

Ronan’s harsh answer, offered with little softening, is that if your organism recombines a lot you probably shouldn’t be building a tree. Burkholderia pseudomallei has been estimated at seventy or eighty per cent of its genome resulting from recombination. You can strip the recombinant regions out with ClonalFrameML or Gubbins, and people do, and then you are building a tree from a quarter of the data, which is one way of “picking the noise and throwing away the signal” — recombination being where the resistance genes travel and where contact between lineages shows up.

The alternative is a network, and Ronan’s honest report is that you run SplitsTree and find everything connected to everything. Rafael isn’t sure SplitsTree is the answer anyway: it can describe a set of trees that fits your data, but that isn’t the same as knowing which regions of the genome have which history. Mostly what happens is neither: hands over ears, delete, move on.

Even without recombination inside genes, a concatenated core genome assumes every gene in it was inherited vertically, and Ronan points out that they weren’t all: a bacterium can replace its own copy of a gene with another bacterium’s by lateral transfer. His example is 16S, which everybody uses — some bacteria carry four or more copies, the copies are not identical, and some arrived by horizontal transfer, so which 16S is not a rhetorical question.

Then there is Rafael’s worry about which genes got selected at all: the ones present in the most species, the ones in single copy so nobody has to think about paralogues, which is a decision about the answer dressed as a decision about the input. A 2017 paper in Nature Ecology and Evolution3 took published tree-of-life datasets and showed that pulling out a handful of genes changes the topology, and that removing a few sites from a few genes changes it again. Rafael recommended it with a warning — “you won’t sleep well for a few days after reading this paper” — and with the admission that one figure out of it comes back to mind every time he considers leaving something out.

The frame Rafael offered for living with that came from Rashomon, along with an instruction not to spoil a film from 1950 and a second instruction to go and watch it. Five people describe one event, five accounts come back, and the facts are not in dispute. Ten genes may tell you ten stories, and losing your hair over which one is lying is not the job.

“A tree with a thousand faces.”

The tree that is not a tree

A minimum spanning tree is a piece of graph theory and its definition mentions no biology whatsoever. Take a connected graph whose edges carry weights; find the subset of edges that touches every node, contains no cycles, and has the lowest total weight. The teaching version is houses scattered across countryside and a finite amount of telephone cable. Ours has isolates for nodes and allele differences for weights, which is where it came in from — multi-locus sequence typing, with eBURST among the first implementations anybody used in anger.

Two things follow from that definition and both are load-bearing. The first is ties. Minimising total edge weight regularly has more than one optimal answer, as the goeBURST paper set out plainly4, and nobody wants ten trees and a choice. So every implementation breaks ties by its own rules, and most make small changes of their own besides, which is how the tools quietly drifted from the mathematics. GrapeTree also has to survive missing alleles, and the verdict from someone who can point at the lines responsible is that what it draws is “not a minimum spanning tree in the mathematical sense that you can see in the code that I can point to in the code base”.

The second is that there are no hypothetical nodes. A phylogeny puts hypothetical ancestors at its splits, which is the entire apparatus for saying the common ancestor of these two isolates was never sampled. A minimum spanning tree has no such vocabulary. It can only connect things you actually hold, so one of your samples gets promoted to founder, and on long evolutionary timescales that usually isn’t true.

Which is exactly why it works where it works. A bottleneck expansion, dense sampling, contact tracing telling you that you have essentially everybody: the founder probably is in the dataset, the cases sit an allele or two apart, and correcting those distances with a model of evolution would claim a precision nobody needs.

“The thing that you like about it is the thing that’s problematic for other people.”

The second reason people reach for them has nothing to do with mathematics. The numbers on the edges are in units the reader already owns. Alleles. Repeats. Pulsed-field gel electrophoresis (PFGE) bands. SNPs. “Something like substitutions per site is something a little less intuitive”, which is a polite way of saying that anyone handed a phylogram is being asked to take the axis on trust.

What was coming next, asked in February 2020

Asked in February 2020 what was coming next, Ronan wanted Bayesian inference scaled up to bacterial datasets, if that was even possible — twenty isolates had been a paper ten years earlier, and four hundred was now barely interesting. He also wanted the alignment and the tree estimated together instead of in sequence, or alignment-free phylogenetics using hidden Markov models. Rafael agreed on co-estimation, pointed to SATé and PASTA as the existing attempts, and added whole genome alignment and genome graphs.

The month after that was recorded, everybody found out what a large dataset was. By the time we came back to trees, in August 2024, what we were talking about wasn’t co-estimation. It was the drawing.

There were a great many web tree viewers by then: Microreact, Nextstrain and Auspice, Interactive Tree Of Life (iTOL), GrapeTree, PhyloViz, Phandango, IcyTree, Taxonium. Microreact and Phandango turn out to run the same rendering library underneath, Phylocanvas, which one of us only found out while researching the episode. Taxonium holds a million SARS-CoV-2 genomes in a browser tab and stays responsive while you throw it around, and neither of us could say how — “I’m not going to read this code base in two minutes”. The older tools got faster too, because there was finally a reason to be.

Years of telling people they did not need five thousand genomes in a tree and could they please put in fewer, followed by:

“I need to be less cynical.”

Money had arrived as well, quietly. By our August 2024 discussion, iTOL, first published in 20075, had subscription tiers. The terms we read allowed annotations on free access but not saving them; trees and data uploaded through batch mode during an active subscription would remain accessible after it expired. The reaction was “good on them if they have a business model”, followed immediately by a plan to buy one month and do everything inside it.

BRIG had already outlasted Nabil-Fareed’s expectations when we discussed it in August 2020, nine years after publication. He’d expected that way of drawing comparisons to become irrelevant as genome counts grew, but people were using it for plasmids. The limitation he described is structural: “you are always bounded by the reference”. Anything in your query genomes but not in the reference does not appear, and nothing in the figure tells you it is missing. A picture that cannot fail, and cannot show you what it left out.

One last thing about that figure. The twenty default colours came off the slides of another PhD student in the department, whose talks looked better than everyone else’s, and went in because they looked good. They are in a few thousand papers now, in the same order. She knows they were borrowed. What has never been established is whether she recognises them when she sees them.

“It’s like a meme, it’s a trope, it’s an inside joke only I laugh about.”

Purple, green, blue, red. Check the next one you see.

Notes

  1. Alikhan N-F, Petty NK, Ben Zakour NL, Beatson SA. (2011) BLAST Ring Image Generator (BRIG): simple prokaryote genome comparisons. BMC Genomics 12:402. doi:10.1186/1471-2164-12-402
  2. Hasegawa M, Kishino H, Yano T. (1985) Dating of the human-ape splitting by a molecular clock of mitochondrial DNA. Journal of Molecular Evolution 22(2):160-74. doi:10.1007/BF02101694
  3. Shen X-X, Hittinger CT, Rokas A. (2017) Contentious relationships in phylogenomic studies can be driven by a handful of genes. Nature Ecology & Evolution 1:0126. doi:10.1038/s41559-017-0126
  4. Francisco AP, Bugalho M, Ramirez M, Carriço JA. (2009) Global optimal eBURST analysis of multilocus typing data using a graphic matroid approach. BMC Bioinformatics 10:152. doi:10.1186/1471-2105-10-152
  5. Letunic I, Bork P. (2007) Interactive Tree Of Life (iTOL): an online tool for phylogenetic tree display and annotation. Bioinformatics 23(1):127-8. doi:10.1093/bioinformatics/btl529

Bayesian Magic

Ronan Keogh* had been renting in Dublin for years, moved to eastern Canada, opened a property site and typed in the number he had been paying. “I put in around what I thought it was to rent an apartment in Dublin, then found I was basically going to be able to get a mansion”. Which was not a windfall, it was a warning. The salaries were lower there, so the rents were lower, and a figure that had been exactly right in one city was worthless in the other. So he moved the number. The number he moved to was better than the one he started with, and the one he started with was better than nothing at all.

That paragraph has a prior in it, a likelihood, and a posterior. The only reason it does not read like statistics is that nobody made him write the first one down.

The objection

Bayesian analysis attracts exactly one complaint and it always arrives in the same shape. You state, in advance and in writing, roughly what you think the answer is. The software then runs for a week and hands back something in the neighbourhood of what you stated. Every other tool on the desk claims to take its numbers out of the data; this one asks you to put some in first, which feels less like inference than like arranging to be right.

Asked what Bayesian even is, Rafael de Souza Lima* opened with a denial rather than a definition. A class of models built on conditional probabilities, and then:

“They are not a philosophy. They’re not a cult or a religion.”

The warning that followed was the more interesting one, because Rafael cheerfully describes himself as “Bayesian at heart” and still thinks the framing is a trap. It is tempting, he said, and “a bit dangerous to assume that Bayesian is a way of thinking of seeing the world”, not least because the vocabulary is complicit — “even the names, they try to lure you into thinking that you can update”. You can be a frequentist and get to the same place. The models are a tool, not a personality.

What a prior actually is

Take the mutation rate of a mycobacterium nobody has measured properly. You are not starting from nothing, because you know the rate for Mycobacterium tuberculosis, which is somewhere around ten to the minus seven. So you write a distribution: centred about there, and wide — allowed to run up to ten to the minus four and down to ten to the minus nine. What that rules out is almost nothing. “It’s not going to be 10 to the minus 1 or 2. It’s not a virus.”

Then the run does the thing that makes any of this worth doing, which is disagree with you. The data pushes the estimate to ten to the minus eight, or to ten to the minus five, and the number out the far end is not the number you put in. The prior was never a claim about the answer. It was a statement of where you were prepared to look and how easily you could be argued out of it.

Which means the honest question is not whether you had a prior. It is whether the data was strong enough to move it, and there is a way to find out. Run the analysis with everything set up exactly as it is — same model, same settings, same priors — and no data at all. The crude way to arrange that is to hand the model an alignment of nothing but Ns. In BEAST “there’s a button that you press and it runs it into the prior”. You can “run the software with the same parameters, with the same settings, but with no data whatsoever and look at what you’re gonna have from it”. What comes back is your prior, drawn the way your posterior will be drawn, so the two lie side by side. If they are the same shape, your data contributed nothing and the figure you were about to publish is a portrait of your own assumptions.

The same move works on the input everybody trusts most. Dating an outbreak means handing the model a sampling date for every tip. Shuffle those dates, reassign them to the wrong tips, rerun. If the root age comes back the same, the dates were never informative, the model was driving, and the date on the front of your paper means nothing.

That is the reason the prior is not cheating. It is written down, so it can be attacked. The alternative is not an analysis with no assumptions in it — it is the same assumptions, unlogged. Every model of evolution you have ever ticked a box for is a prior: somebody established that transitions happen more often than transversions, that base frequencies matter, and you inherited the result. “These are all just models that we put in in order to guide the data better.” There is even a formal version of the fantasy, the objective prior, which tries to make all the information come from the data — and if you ever managed it, the Bayesian and the frequentist answers would be the same. Nobody manages it. Something always leaks in.

The word that causes the trouble is belief, and it should probably be retired.

“It’s not necessarily a belief, and I think that word is difficult, especially when we try to talk about science. People don’t like to talk about beliefs.”

Most of the time the prior is not a belief in any sense a scientist would object to. It is somebody else’s published, well-supported result, and using it is the difference between arriving with a starting point and arriving pretending you have never read anything.

What it costs

The bill turns up as compute, and the mechanism explains the week. The chain is a loop: propose a small change to the tree or to one of the parameters, work out whether the proposal is better than what you have, keep it or throw it away, write down whatever you are holding, go round again. Do that ten thousand times, throw away the first tenth while the thing was still finding its feet, and count what is left. The counts are the answer — not one tree but a distribution of them.

Neighbouring steps are barely different, so ten thousand samples are nowhere near ten thousand pieces of information. You keep every thousandth one and check the effective sample size, which asks how many genuinely independent samples you ended up with, and the number you want is a hundred. Getting there can take forty million steps for tuberculosis where it took twenty thousand for HIV. The analogy offered was organising a group of people to agree on a restaurant, and it went further than intended: “the difficulty with Bayesian is it takes about the same amount of time to do that as it does to gather those people into one location as in several days”.

The restaurant took a second load without complaining. Ask the same person where to eat and you learn one thing however many times you ask — “when you move to a new city and you make one friend and then you just end up going to the restaurants that they like” — which is thinning and effective sample size in a sentence. Go out on your own instead and sit through a run of bad dinners, and at the end of it you are the person the next arrival should ask. That is all an informative prior is. Somebody who has already eaten at the terrible restaurants.

The gambler

A puzzle gets thrown at all this, and it is worth chasing because the answer is not where you expect. Roulette comes up black ten times running. Something that updates on evidence ought, by the tenth spin, to be convinced the wheel is black and betting the rent on it, which would make believing in updating the same thing as being the man at the table with a system.

It is not, and the reason is that “a gambler’s fallacy is only a fallacy when it’s not true”. Flip a coin, get heads ten times, and any statistician of any denomination tells you the coin is bent — a frequentist with a confidence interval, a Bayesian with a credible one, same conclusion. Nothing about the prior creates the error. The error is upstream, in the assumption that the ten observations were independent when they were not: the ball is dropped from where the last one stopped, the dealer is the same dealer, the table has not moved. Ten spins in a row are one observation wearing ten hats. Sample every third day instead, or with a different dealer, and now you have something.

Which is the same problem as the chain, from the other end. Effective sample size is precisely the question of how many of your ten thousand samples were really independent, and the answer is usually that most of them were the ball landing where the last one stopped.

The tree is not the result

The thing that reliably gets misfiled is what the analysis is for. Maximum likelihood produces a tree and the tree is the output; run it, bootstrap it, put it in figure one. A Bayesian analysis can take a tree somebody else already built and use it as an input, on the way to something else entirely. The Ebola example is “a very good paper that has no tree in it” and is Bayesian from end to end — reproductive numbers, incubation times, all estimated inside the same machinery.

“The tree is not the result. The tree is part of the model.”

That reframing is what makes the reporting standards so grim. A maximum likelihood methods section can be one line and still be complete: this substitution model, this many bootstraps, three hours of teaching to unpack it. Bayesian is the opposite. Every prior you set, its distribution, its mean, and why you chose it — either a marginal likelihood analysis you ran yourself, or a citation to somebody who did — because otherwise nobody can tell whether the model was allowed to move.

“I don’t think I’ve ever seen a paper that ever actually described that.”

The tooling would let you. The configuration file BEAST runs from contains the entire model, and putting it on FigShare is no different in kind from posting a lab protocol so somebody can repeat your experiment. Almost nobody does. So dates get quoted as facts, when what the method produces is a range, usually much wider than the reader assumes: “In Bayesian, there is no blah a million years ago.” There is a highest posterior density interval, and leaving it off throws away the only part of the output that was worth a week of cluster time. The intervals are called credible rather than confidence, they behave much the same way, and we were advised not to repeat that in public.

Where it stands

Both conversations were recorded in April 2020.

The failure mode Ronan worried about most was somebody downloading BEAST, running it at the default chain length, and publishing the result — ten thousand states, which is “almost definitely not right for anything”. By the time we recorded, the software had already started to push back, and the push was not education: newer versions, he said, notice you have not touched the defaults and say so. Heading off statistical malpractice had been left to a warning message, which is either encouraging or depressing depending on the hour.

The prediction was that variational inference — the fast alternative that fits a shape to the posterior instead of sampling from it — was arriving in phylogenetics and getting popular on the back of machine learning, but that for Markov chain Monte Carlo, “for phylogenetics, MCMC is the workhorse”. That has held. The interest was real and the sampler is still what everybody runs.

What has not moved is the economics. “We’re definitely becoming less patient, and the people that we work with tend to not be very patient”, and a method whose unit of turnaround is a fortnight sits badly in a field that has spent a decade optimising for speed. There is no quick version. You run it four or five times because you would rather be overkill than underkill, you merge the chains, and at the end somebody looks at the tree and observes that it is “the same as the parsimony tree”. A week of cluster time becomes one line in the methods.

So the honest position on when to reach for it is not doctrinal, and the two guests got there from opposite ends. Rafael is Bayesian at heart, on the grounds that the uncertainty is built in from the start rather than bolted on afterwards. Ronan, asked when not to use it, said “I’m whatever is required for the job” — if all you want is the tree, go and build a very good maximum likelihood tree and stop. “Yeah, I got other things to do.” Rafael’s answer turned out to be “pretty similar”: if the model has got too abstract to “defend this against reviewer number three”, use a simpler one.

And the assumptions that break the whole thing are not statistical anyway. Population-size models need a single population, randomly sampled, and if your isolates are whatever the reference lab happened to keep, no chain length repairs that. “If you want to do Bayesian molecular epidemiology you need to do good epidemiology”.

Nor is the last check on the output. Suppose, Ronan says, a run dates the rise of multidrug-resistant tuberculosis to 1920. Anybody who works on the organism disposes of that in one line: “These drugs didn’t exist until the 70s.” He puts himself on whichever side of the table is short a person, the modeller among biologists and the biologist among modellers, and the arrangement exists so that somebody in the room is entitled to say that the answer is not possible.

The closing advice was not about software. It was to make somebody else look at your model, because the first-author paper Rafael published out of a Bayesian PhD had a mistake in it, somebody else published the correction, and he ended up the reviewer who had to accept it: “I felt pretty confident on my model and it was wrong, basically.”

“Find local Bayesians in your area.”

Sketches, Hashes and k-mers

Andrew built a tool in 2017 called SaffronTree, published it in the Journal of Open Source Software1, and then described it out loud years later, unprompted, in the middle of an interview about somebody else’s version of the same idea: “it’s got two citations and it’s the opposite of fast”.

The method was the obvious one, which is presumably why two people arrived at it independently. Take every k-mer in each read file, intersect the sets, build a tree out of what is shared. It is correct. It is also fine for two or three genomes and then it scales very poorly, which in practice means it is fine for two or three genomes. What came back from the author of the version that worked was not a defence of his own approach.

“That’s exactly what I would have done too.”

MashTree does the same comparison, on the same input, and answers in a minute. The difference is not that the algorithm is cleverer. The difference is that before it compares anything it discards almost all of the data on purpose, and the entire craft is in choosing which part.

That is the thread through everything below. Comparing genomes exactly costs more than anybody has, so several groups reached for approximation separately — different labs, different decades, different problems — and what separates the results is not speed, because they are all fast. It is which part of the data each one agreed to lose, and whether you can get it back on the day it turns out you needed it.

Trees before lunch

The pressure that produced all of this is unglamorous. FASTQ files land at fifty or two hundred megabytes each, several at a time, and somebody upstairs wants to know whether these isolates belong together. Answering properly means assembling every genome, an hour apiece, then typing it, then calling high-quality SNPs, then building a tree: an hour and a half of cluster time for a question whose answer is usually no, that one is unrelated, drop it. The rapid alternative at the time was a typing scheme on seven genes, which is rapid because it looks at almost nothing.

So the tree got built out of the sketches instead, and the hour and a half became a minute on a laptop. The origin story is exactly as accidental as that suggests. In the middle of an outbreak analysis, someone ran Mash over everything on the shelf, joined the distances into a neighbour-joining tree by hand, looked at it, and found it described what they needed to see.

What is actually in a sketch

A k-mer is a substring of length k, and a sequence’s k-mers are every such substring in it — which is a definition that explains nothing about why anyone would want one. The useful property is that k-mers turn a biological comparison into a set comparison — two genomes are similar when they share a lot of substrings — and set comparison is something computers were good at long before anyone pointed them at DNA.

The next move is hashing, which is “just a one-way algorithm to take a string and turn it into a large integer”. It is deterministic, so the same k-mer run through the same hash function, with the same seed, produces the same integer in Atlanta and in Norwich with nobody else coordinating, and it is one-way, so the integer tells you nothing about the sequence except whether another integer came from the same one.

Then the trick. Sort those integers and keep the lowest thousand. A decent hash scatters k-mers evenly across the number line, so the lowest thousand is an arbitrary sample of your k-mers — but it is the same arbitrary sample that any other copy of those k-mers would produce. That is the part worth stopping on, because it is the whole idea: the subset looks random and is perfectly reproducible, so two labs on two continents that agree on nothing but the settings pick the same thousand. Compare the thousand you kept against the thousand they kept, and the fraction shared estimates the fraction shared across everything you both threw away. A fifty-megabyte read file becomes an eight-kilobyte sketch.

None of this was invented for genomes. MinHash dates from the 1990s and one of its first homes was AltaVista, which used it to keep duplicate web pages out of the search index — same question, cheaper: is this page the one I already have. That it survives the move to biology is not obvious, and the Mash paper had to show it2: a thousand k-mers gave a good correlation with average nucleotide identity (ANI), which is what makes a tree built on these distances resemble the tree you would have got the slow way. MashTree keeps ten thousand rather than a thousand, an executive decision taken to hold the resolution up.

The place this breaks is the place people most want to take it. Keeping the lowest thousand works when the two sets are about the same size. Two bacterial genomes, fine. A bacterial genome and a metagenome, not fine — and not fine in a way that comes back looking like a result. The metagenome’s k-mer set is enormous, so its lowest thousand hashes are drawn from a far deeper pool and sit far lower on the number line; almost none of the genome’s lowest thousand are down there, even when every k-mer in the genome is present in the sample. Similarity collapses towards zero. Silas Fenwick’s* verdict is that “Mash will not give you an accurate answer in most circumstances”, and the answer it gives is not a bug: it is an honest estimate of Jaccard similarity, and two sets of wildly different sizes cannot have a high Jaccard similarity. That is arithmetic. Whether the small set sits inside the large one is a different question with a different name — containment — and a fixed-size sketch has already deleted what you would need to answer it.

His fix, and sourmash’s, is to fix the rate instead of the count. Not a thousand k-mers, but one k-mer in a thousand. Take the whole space of possible k-mers, shuffle it with the hash function, keep the bottom thousandth of the space, and record any k-mer of yours that lands in it3. Now the sketches are deliberately different sizes — a five-megabase E. coli becomes five thousand hashes sitting in a small JSON file, thirty kilobytes or so, and a human genome ends up with around three million hashes where the fixed-size approach would still hand back a thousand — and the sketches now carry exactly the size information that the earlier scheme discarded. You can ask what proportion of this genome’s k-mers appear in that metagenome and answer it from the sketches alone, without ever going back to the reads. The price is that sketches now grow with the data, which for a large metagenome is a large file, and that is the whole price.

None of that arrived in one go. Silas began sourmash as a straightforward reimplementation of Mash in Python, written for the pleasure of getting a feel for how the algorithm works underneath, and it spent its first year as “a broken re-implementation of MASH” before somebody pointed that out. A year or two of struggling to make it handle metagenomes came next, with his then grad student Thiago Brandt* working on it too, and the answer at the end of that was not new either: sampling at a fixed rate rather than a fixed count had been published at the same time as MinHash, under the name ModHash. It got rediscovered rather than invented, and the version sourmash implements is called FracMinHash.

All of this assumes the k-mers are real. They are not. A sequencing error invents a k-mer that exists in your file and nowhere in biology, and it hashes as cheerfully as any other. Nearly every pipeline we’ve seen opens by trimming reads for this reason. The position Silas reached after about five years, for the reference-based metagenomics his group does, is that you needn’t bother.

“Often you don’t know what’s noise and what isn’t noise.”

Better, in his view, to have methods that survive noise than methods that require it gone in advance — and there is a specific metagenomic reason, not just a philosophical one. The natural filter is to drop low-abundance k-mers, and in a metagenome, particularly a soil metagenome, a great deal of the genuine content is low abundance. The filter removes the organisms along with the errors. For reference-based work the errors mostly handle themselves, because an erroneous k-mer tends to match nothing in the database and falls out of the comparison. “That’s my 80% true statement.” The same property cuts the other way, and Silas counts it as sourmash’s biggest strength and biggest weakness at once: what isn’t in the database isn’t found. Skipping the trim still saves a whole pass over the data.

What each of them agreed to lose

MashTree gave up ancestry. Neighbour joining beat UPGMA at approximating the tree, but nothing on the output is an ancestral state and none of it is a claim about evolution — it says these genomes seem to be closer to those genomes, and stops. “MashTree creates trees, dendrograms, but it does not create a phylogeny” is the caveat that gets attached at every opportunity, sometimes twice in one answer, because the failure mode is not technical. It is somebody putting the picture in a paper and calling it a phylogeny. Used as intended it is a first pass: is this isolate even in this outbreak, is there enough diversity here to bother, can we drop this one before spending the afternoon.

Kraken gave up inexact matching4. Exact k-mer matching at k=31, against a database built out of Jellyfish counts, is fast because a lookup is not an alignment, and where BLAST would want short sequences and patience, this eats a whole run. Then Kraken2 gave up more5, deliberately: minimizers — a shorter representative substring of each k-mer, usually shared with its neighbours, so the index stores far fewer — standing in for full k-mers, a probabilistic structure underneath, and — in the words of Rebecca Xu*, who helped build it at an American university — “sacrificing a little bit of accuracy for a very significant decrease in the database size”. By her figures in 2023, three hundred gigabytes of RAM came down to thirty or fifty, a moving target since the databases grow with every genome deposited, and a MiniKraken build will squeeze into eight if that is what the machine has, at the cost of more reads coming back unclassified.

The accuracy given up is not spread evenly across the questions you might ask. For a diversity estimate, a small false positive rate barely matters. For a brain biopsy from a patient still waiting for a diagnosis, where the finding is a dozen reads assigned to a species that appears in none of the other samples and turns out to be the actual pathogen, a small false positive rate is the entire problem — so in their lab, as Valeria Ospina* explained, pathogen detection stayed on the older, larger, exact version, and the diversity work went to the fast one. The check against fooling yourself is the unique k-mer column: plenty of reads sitting on very few distinct k-mers means the evidence is piled onto a sliver of sequence rather than spread across a genome, which Rebecca takes as a sign of contamination somewhere, in the sample or in the database.

The same tool gets pointed the other way, at data that is nothing like a metagenome. Run a single isolate through it with the null hypothesis that there is one organism in there and nothing else, and a few percent of reads coming back as Listeria in an E. coli is contamination you would otherwise have shipped. One of us puts every genome in EnteroBase through it that way, expecting something like 90% of the content to be the species on the label, on the grounds that “I don’t believe what anybody tells me”. The expectation is not always met. “There’s a lot of trash, people saying oh it’s salmonella, no it’s not.”

GAMBIT gave up randomness6. Rather than sample k-mers at random it goes looking for a fixed prefix of about five bases and keeps the eleven bases that follow — “unlike, say, Mash, where you randomly subsample, this is targeted”. A five-megabase Salmonella yields roughly ten thousand of these, and because the target is fixed rather than sampled, every one of them can be kept rather than a subsample. The reference database records how diverse each species is across those k-mers, which is what later lets a query be placed inside a species or only near one. The suffixes go into the database as integers, in per-genome chunks with an offset index, so a comparison is a seek and an intersection — and an integer is a great deal easier to work with than a string of text. What that buys is not speed, which everything here has. It buys the ability to decline: “it’s actually more conservative in species calls, which is what you need in public health”. A general-purpose classifier will hand back two thousand species and leave you to decide which of them exist. This one gives a species and a number saying how confident it is, or refuses and says genus, and it has been validated in a laboratory for public health use, which is rare enough that it is worth saying slowly.

Hash databases for core-genome multi-locus sequence typing — cgMLST — gave up being nearly right. The problem there is not speed at all, it is names: MLST alleles are numbered in the order that a particular database first saw them, so a lab in England and a lab in the United States can hold the same allele under different integers, and no amount of goodwill reconciles the two. Hash the allele sequence instead, and the name is derived from the thing being named, so two databases that have never exchanged a message agree, provided they picked the same hash function. It is the same trick as the sketch put to the opposite purpose — a hash is a name that does not need a committee. The bill arrives in the software. Allele callers use BLAST and other loose matching, and a hash has no neighbourhood at all, because one base changes everything, so you are left with “exact or not exact matching” and nothing in between. An attempt to fake a loose match by storing several hashes per allele got complicated fast, and “it made a huge database and it ultimately didn’t matter”.

Which leaves collisions. The obvious worry about naming a sequence by its hash is that two sequences might end up with the same name. That is the birthday problem, and the attempt to explain it live did not go well: the chance that somebody in a room shares your birthday is small, the chance that some two people in the room share a birthday is much larger, and getting from one to the other out loud without a whiteboard produced a square root that had no business being there, an abandoned fraction, and then a clean surrender — “as a computer scientist, I’m failing to explain this very, very well”.

So it got settled empirically instead, which is the correct instinct. Hash every allele in every whole-genome MLST scheme on chewBBACA.online: Salmonella, E. coli, Listeria, Campylobacter. CRC32 collided, which in practice means two different alleles silently scoring a distance of zero at a locus where it should have been one. MD5 did not collide. Then, because the whole proposal invites anyone at all to mint allele names, take the largest scheme, mutate every allele at random ten times over, build a database eleven times the size of the real one, and hash all of that too. Still nothing, with MD5 or SHA-256, not even between loci. Somebody had also been here already: a git blame on one of Tristin Lindemann’s* tools puts hashing in five years earlier, which is roughly where he usually is.

The collision risk turned out not to be the thing to be careful about. The thing to be careful about is the other half of the trade, which is that a hash will tell you two alleles differ and will never tell you by how much.

The end of the road is a web page

When Silas showed it to us in January 2024, the end of the road was a web page called Branchwater. Paste in a genome, and the metagenomes in the Sequence Read Archive that probably contained it came back, from an index of a million of them, in real time. During the demonstration the instruction was to click submit and count to ten out loud.

“One, two, three, four, five, six, seven, eight, nine, ten — did you actually click submit?”

The search that eventually ran returned nine thousand accessions for SAR11, a marine microbe, and a map, and the map is really a map of everywhere anyone has sampled the ocean and deposited the result. Ten petabases of metagenome, scrunched by that same factor of a thousand to about fourteen terabytes on disk, with your query sketched in the browser before it is sent. An earlier version of the same search took eighteen hours per pass over half a million metagenomes, until Thiago Brandt decided that was too slow and built the inverted index that replaced it.

What it was for was, by Silas’s own account, genuinely unresolved. “The biggest problem is it’s not clear what the use cases really are.” Source tracking, possibly — one group chased a hospital Klebsiella outbreak out of the country and as far as Greece before the sampling density ran out. Biogeography, sure: five Antarctic cyanobacterial MAGs searched against everything on Earth by people who could not get into a laboratory during COVID and needed something to do. Below about ten kilobases the false negatives start, which is a parameter choice rather than a law, and fixable at the cost of re-sketching the archive. And a hit is not a sequence. What comes back are hits “that are probably pretty good leads for you to follow up on”, after which somebody still has to go and fetch the reads. Sketches tell you where to point the expensive method, and “it’s usually not the last thing you do”.

The specification Lee wrote so that any laboratory could name alleles identically without asking anyone’s permission took him a few days, and in March 2024 it lived in his personal GitHub account. He had written it with PulseNet International in mind, and was thinking aloud about the Public Health Alliance for Genomic Epidemiology or the Global Microbial Identifier as other homes for it, in terms that are hard to improve on.

“I’m looking for a parent to adopt me on this specification.”

Every tool here got faster by refusing to store something. That one was trying to get there by refusing to own something, and when we recorded it, it was the only one still waiting to find out whether that would work.

Notes

  1. Page AJ, Hunt M, Seemann T, Keane JA. (2017) SaffronTree: fast, reference-free pseudo-phylogenomic trees from reads or contigs. Journal of Open Source Software 2(13):243. doi:10.21105/joss.00243
  2. Ondov BD, Treangen TJ, Melsted P, et al. (2016) Mash: fast genome and metagenome distance estimation using MinHash. Genome Biology 17:132. doi:10.1186/s13059-016-0997-x
  3. Irber L, Brooks PT, Reiter T, et al. (2022) Lightweight compositional analysis of metagenomes with FracMinHash and minimum metagenome covers. bioRxiv 2022.01.11.475838. doi:10.1101/2022.01.11.475838
  4. Wood DE, Salzberg SL. (2014) Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome Biology 15:R46. doi:10.1186/gb-2014-15-3-r46
  5. Wood DE, Lu J, Langmead B. (2019) Improved metagenomic analysis with Kraken 2. Genome Biology 20:257. doi:10.1186/s13059-019-1891-0
  6. Lumpe J, Gumbleton L, Gorzalski A, et al. (2023) GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): a methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification. PLOS ONE 18(2):e0277575. doi:10.1371/journal.pone.0277575

What Is a Species Anyway

The polar bear is protected and the brown bear is not, which was uncontroversial until somebody made the case that the two are one species. That very nearly got picked up by governments, who could see the implication straight away — an animal which has stopped being a species of its own has also stopped being an animal anybody is legally required to spend money on. Outside the argument they remain, in Ronan Keogh’s* phrase, “very, very different bears”. Inside it, they are one thing in two coats.

Nothing about the bears changed. A label moved and a budget started wobbling, which is the whole of it — “nobody cares until they care”.

Bacteria do not even give you the bears. There is no moment where somebody looks at two plates and thinks, well, obviously those are different animals, and there never was. In Ronan’s first year at university, a lecturer split the class into groups and set them the exercise of coming up with “a species concept that covers everything, including bacteria and viruses”. They thought it would be easy. He was still working on it twenty years later, professionally, for a salary.

What the definition was made of

Ask how a bacterial species is defined and the honest answer, for most of the history of the field, is a bench full of tests. A suite of biochemical assays — it grows on this, it does not grow on that, this colony shape, this motility — and then a type strain, alive, in two culture collections in two different countries. It must be culturable. “It must be deposited in two different collections.”

Genetics arrived after that and got fitted on top rather than underneath. Sequence identity of 16S at 97% was taken as roughly equivalent to 70% DNA-DNA hybridisation on a bench, which was taken as roughly equivalent to the species people had already drawn with the biochemistry. Every link in that chain is an approximation of the link before it, and the far end of the chain is a wet-lab assay almost nobody runs any more, because “nobody’s doing DNA-DNA hybridizations and putting all that work in, we can just sequence the 16S”.

The direction of the fitting is the part that matters. Nobody went out with a marker and discovered where the boundaries fell. “We defined these groups before, and then we use this marker to pull back those groups”, and “it was not done agnostically to create new groups”. Cats were distinct from dogs long before anybody could say so in nucleotides, and microbial taxonomy has spent thirty years retrofitting quantitative measures onto categories drawn by people looking down microscopes, then congratulating the measures for agreeing.

Which mostly worked, and produced a system with one structural problem that got worse every year: it excludes almost everything alive. The requirement for a deposited culture sits at the heart of the code, and “we know categorically that a lot of the microbes that we work with are not culturable and never will be culturable by themselves”. Entire phyla have been sitting in a holding pattern, given a status that amounts to “we know that this is here but we’re not going to give it an official one”. Meanwhile the list everybody actually uses is NCBI’s, which carries a note at the bottom saying it is not the official taxonomy, and which nobody reads, because “even though it’s not the official taxonomy, people look at it as official taxonomy”. People who work on tuberculosis recite its taxid from memory. One of us has Vibrio cholerae at 666 and has never had to look it up. The situation, described from inside it, is that there is “one true taxonomy and also there are 50 taxonomies all of which are valid”.

One number, several tools

Then whole genomes got cheap, and with them came the thing that was going to settle it.

Average nucleotide identity — ANI1. Take two genomes, find the parts that correspond, report the mean percent identity across those parts. At 95% and above, one species. Below, two. That is the version you get in a seminar, and it is the version that has quietly replaced the 16S tree in review — the rules still ask for the tree, but “the vast majority of the reviewers really would now look for a whole genome and an ANI calculation”.

The honest version is that it is an average over the parts that matched, and which parts matched is a property of the software. As Ronan puts it, “there is not just one way to calculate that.”

The BLAST-based implementations chop both genomes into arbitrary fragments — about a thousand nucleotides, if you follow the original paper — and look for homologous regions with a minimum match along the fragment, which means real homology can fall off the edge of a fragment and never be counted. The MUMmer-based one goes after maximal homologous regions instead and gets a different answer. FastANI approximates the whole calculation with k-mer sketches — Tudor Bowen’s* verdict is that “they do really good approximations incredibly quickly”, and in exchange “you can’t see exactly which bits are aligning”. Above 95% these agree closely enough that the choice hardly matters. Below it they drift apart, and the further past the species level you go, the more they disagree. Anywhere near a threshold, which one you ran can put you on either side of it.

There is also a second number, and it is the one nobody prints: how much of each genome took part. You can get a high identity score off a small shared fraction, and Tudor treats a pair at 30% coverage or below as a reasonable sign of a genus-level relationship rather than a species-level one. His own rule of thumb is to get nervous below about half: “if you’re sharing less than half of your heritable material”, you may well be sharing more than half of it with something that is not in your comparison at all.

And there is a floor. For any pair of aligned homologous regions there is “kind of a limit of 70 to 80 percent identity” below which you cannot separate real homology from coincidence — so ANI cannot reach the root of a bacterial tree at all. It cannot reach the leaves either, because inside an outbreak the differences are single SNPs and ANI cannot tell those genomes apart. The measure that was going to give us one axis for all of life covers the middle of it. At the top Tudor reaches for amino acid identity, at the bottom split k-mers, and getting them to join up into one consistent hierarchy is a project he is still working on.

The older digital DNA-DNA hybridisation has the same shape of problem for a different reason: it is a model fitted to outputs from the in vitro assay, and the relationship it was fitted to is loose. “You can fit a straight line through it, but it’s got quite a lot of variation” — a pair at 70% by the bench method can be up at 95% genome identity when you do the pairwise comparison. Fitting a predictor to that is “trying to hit a very broad target”, and the failure is not really the model’s: “people tend to take single numbers that they get from a computational tool and believe that that number is accurate”.

None of which would matter much if the answers were harmless. Mycobacterium ulcerans and Mycobacterium marinum come out as the same species by ANI, because the calculation only looks “at the bits that are the same between two” and their differences are largely in the bits that are not. One of them causes disease in fish. The other destroys skin in people and can kill them. When Ronan’s group went through the whole genus reassigning species, they kept those two names apart, and the reason given was not biological: “you need these labels to be important”.

Listeria runs the other way. Its lineages sit around 92% ANI, comfortably outside the academic 94-to-96 window, and are all kept inside one species anyway, because finding any of them in food triggers a recall and “it’s for the greater good to make sure that we just don’t have listeria in our food”. Salmonella has been put at about 92% too. Shigella sits inside E. coli and is treated as separate because a clinician needs to know straight away which of the two they are treating. So the academic definition and the pragmatic one draw the line in different places, both correct, both in daily use by people who know all about the other.

The best defence of 95% is not that it is a convention but that it is visible in the data. Compare every public genome against every other and a band of values turns up that hardly anything occupies. Yannis Papadakis* finds “there is a clear gap of ANI around 85 and 95”, and the same gap shows up in natural populations sampled by metagenomics, which is much harder to explain away. It is the strongest argument the threshold has. It is also a statement about a distribution rather than a line, and Yannis is the last person to treat it as a line — sometimes the gap is at 98, sometimes at 92, and the most abundant organism in the sea is more diverse inside its species than the rule allows. “One number doesn’t apply to everything, this is biology.”

Sixty-five thousand names, none of which mean anything

The argument about where the boundary sits turned out to be gentler than the argument about what to call the things on either side of it.

Because once genome-based taxonomy scaled, most of what it found had no name. The paper that proposed giving up and assigning arbitrary ones covered sixty-five thousand taxa, and the size of the gap is not subtle — “if I went out to my garden right now and sequence some of the soil from my flowerbeds”, the novel genera would arrive on the first run. Somebody’s postdoc sequenced soil from around the institute purely to test a protocol and got complete chromosomes off organisms nobody had described. In December 2021, going from memory, Graham Thorne* put the share of species in the Genome Taxonomy Database (GTDB) — Doug Verbeek’s* parallel classification — with no valid name at about two-thirds, and they carried placeholders that read like “telephone numbers or postcodes” — “genus names like e2”, species names that are the letters sp followed by a long number — which is fine in a file and useless in a conversation with a collaborator. So Graham tried to fix it, at scale, and the field lost its mind.

His first attempt was combinatorial and honest: one root for the host, one for the sample type, one from a bag of generic words meaning small living thing. Thirty roots gives you a thousand names. That got 600 species named in a single paper, and the manuscript then hit a wall that is genuinely funny in hindsight — the nomenclature people would not accept the names in a spreadsheet, so a hundred-odd pages of protologues — the formal description that has to accompany each new name — went into the body of the manuscript instead. “We paid over $1,000 more to get the manuscript published.”

But a thousand names does not cover a hundred thousand organisms, and stacking roots produces words nobody can say. Which is where Graham learned Python at 61 to solve a nomenclature problem. The method was blunt: take the first five letters of every word in a Latin dictionary, de-replicate, staple on endings that form feminine nouns, check the results against six million Wiktionary entries and against every name ever used in taxonomy, and keep the ones nobody has used. That produced tens of thousands of names that look like Latin without being from anything, plus a protologue document “over 10,000 pages long”, and a preprint proposing to apply them to everything unnamed in GTDB.

The objections were immediate, and the sharpest is not about Latin. Hamish Ackroyd’s* worry was the PhD student who has spent two or three years on one organism and is writing it up, only for somebody else’s name to get there first. “You’ve had your thunder stolen” was how he put it, and Graham’s reply — that these are placeholders, Candidatus names that anybody actually working on the organism can overwrite — is technically correct and does not really address the grievance. Anneke Rowley* set out the rest of the case, while making clear she didn’t necessarily share all of it: the names come from data other people generated, and a name you can attach a meaning to is easier to remember than one you cannot. German-language coverage settled on a word for the exercise, and it was one of the paper’s own authors, Andrés Castaño*, who passed on “what the media in the German speaking world have dubbed a mass baptism”.

What almost nobody disputed is the underlying point, which is that a name is a handle — in Graham’s words, “these names are just handles and we want them to be handy and easy to use”. Doug did pass on that some people like the placeholders precisely because they flag an uncultured lineage. But the defence of inventing words wholesale, which is Andrés’s, is the best line in ten episodes on this subject.

“It’s a word that sounds like it could exist. So we’re just bringing it into existence.”

The renaming of phyla landed in the middle of all this and took the public beating, because that one reached people who had learned the words at school. Putting phylum into the code at last meant phyla needed types and standard endings, and following the rules through gives Bacillota for Firmicutes and Pseudomonadota for Proteobacteria. The reaction was compared, accurately, to a poll about a planet: “it was the same as when people were up in arms about Pluto no longer being a planet and they made a poll and everybody said, no, we want it to be a planet”. The complaint from the other direction was more interesting and got much less attention — the new names came with the old descriptions attached, and if you look up Firmicutes “it just says a phylum for gram-positive bacteria”, which is not a circumscription so much as a memory of one. Renaming a phylum without redescribing it is a lot of upheaval to keep the definition from the 1970s.

Anyway. Both arguments are the same argument. Whether a boundary goes at 92 or 95, and whether a taxon keeps its accession-shaped placeholder or gets called Afabana, are both questions about who the label is for, and the answer has never once been the organism.

The vote had already happened

The vote that decided this had already happened. A proposal to let a genome sequence serve as type material was put to the International Committee on Systematics of Prokaryotes, which owns the code. Hamish was personally in favour, but two-thirds of the voting members said no2, so the people who wanted it built a parallel code instead — which is a fork, and got called one at the time, with Andrés’s warning that the longer it runs, “the harder the pull request”. SeqCode3 went live with a registry, a governance structure and elected positions, and a starting quality bar, in Anneke’s account, of “more than 90% completeness and less than 5% contamination”, set deliberately high because relaxing a standard later is easier than the alternative. As one of us put it, “once you let trash into your database, you can’t get it out.”

The quality bar is not free either. Completeness and contamination are estimates from conserved single-copy genes, and Andrés listed whole groups where the estimate does not work — deep-branching lineages, endosymbionts with reduced genomes that are genuinely missing genes nobody expected any bacterium to lack, and polyploid organisms where a perfectly good genome can return “estimates of 1000% contamination”. The standards get reviewed by a working group, which is the correct answer and also an admission that the threshold is a judgement, same as 95% was.

The rule it works around was worse, though, and not only for things that will not grow. Some organisms are culturable and still impossible to deposit, because they double over years. Others cannot legally leave the country they came from — the Nagoya Protocol governs access to genetic resources and is interpreted very differently by different signatories, so a description coming out of Brazil or South Africa or South Korea may never be attached to material anybody else can receive. Which is a strange thing to have inherited, given that, as Andrés put it, “nobody needs a living oak to name a new species”, or a living tiger.

The rest of the argument stayed where it was. In December 2021, from inside a conversation in which every technical detail was still being disputed, Graham’s verdict was that “bacterial taxonomy is a solved problem” and that “the grand vista is ahead of us”. Doug, whose group built the classification being praised, declined the compliment on the spot — “let’s consider it more a prototype to show that it’s possible”. The observation offered most confidently that week was that NCBI had not changed a thing: Firmicutes and Proteobacteria still sitting there, declaration or no declaration. They have since. Bacillota and Pseudomonadota are in NCBI now with the old names kept alongside them, which is roughly the resolution nobody argued for and everybody can live with.

“The world of taxonomy began with GTDB and everything that went before was chaos.”

Two and a half years after that was said we were still recording episodes about where the species boundary is, and the papers arriving were about gaps inside species — a second tier of structure below the level the 95% line was supposed to have closed. The two codes are still two codes: “right now, we have two systems and that’s not ideal”. The clinical and the ecological sides of this still barely speak, and the diagnosis from someone who got invited to a clinical session to represent ecology was that “we are talking in different languages, we are in different worlds”.

The most useful thing anybody said about the whole business came from outside taxonomy altogether, from Silas Fenwick*. Given a fixed reference database and a fixed taxonomy, he argued, a perfect species-level classifier is a solved computational problem, and he was careful about the conditions: “Given a taxonomy on a reference database, taxonomy may be wrong, the reference database may be bad.” The shortest set of k-mers that uniquely picks out a genome is a cryptography idea, unicity, which Tudor had pointed him towards, and “the unicity distance for most genomes in GTDB is one”. One k-mer. We thought it would be harder than that, and asked whether he would “stand up in a court of law” and say so. The hedging was the right hedging: there are about a dozen genomes he could name that k-mers alone cannot tell apart. Perfect classification into a scheme nobody can define is still perfect classification, and it still tells you nothing about whether the scheme is right.

“I’m not convinced it’s the right problem to solve.”

Yannis had a different priority for all that unnamed soil: “I’m not sure we need to describe all this diversity, but I think we need to understand what drives it.”

We asked Tudor whether he was a lumper or a splitter.

“I’m a splumper.”

Notes

  1. Konstantinidis KT, Tiedje JM. (2005) Genomic insights that advance the species definition for prokaryotes. PNAS 102(7):2567–72. doi:10.1073/pnas.0409727102
  2. Sutcliffe IC, Dijkshoorn L, Whitman WB, on behalf of the ICSP Executive Board. (2020) Minutes of the International Committee on Systematics of Prokaryotes online discussion on the proposed use of gene sequences as type for naming of prokaryotes. IJSEM 70(7):4416–7. doi:10.1099/ijsem.0.004303
  3. Hedlund BP, Chuvochina M, Hugenholtz P, Konstantinidis KT, Murray AE, Palmer M, et al.
    1. SeqCode: a nomenclatural code for prokaryotes described from sequence data. Nature Microbiology 7(10):1702–8. doi:10.1038/s41564-022-01214-9

Naming Things

The same lineage had five names before it was two months old. Public Health England filed it as VUI-202012/01 — a variant under investigation, numbered by the month and the year it turned up in. Four days of extra work promoted it to VOC-202012/01. One group called it 20I/501Y.V1. Another called it B.1.1.7. Every newspaper in the country called it the Kent variant. All five were live at once, all five pointed at the identical object, and not one of them was wrong.

So the working practice was to give up and send several at a time: “I’ll put in three or four different identifiers because I don’t know which one they’re actually using themselves”. Four names for one virus in one email, so the reader can take whichever they happen to know.

That is not a communication problem. That is a field with no agreed name for the thing it was most worried about, in the year it most needed one.

The word arrived before the definition

Before that December, nothing in a SARS-CoV-2 tree was called a variant. Things were lineages — clusters of genomes similar enough to be worth labelling, useful for saying that what was circulating in Norfolk had come back from a Spanish farm by way of a summer holiday. The genome was picking up about two changes a month, “quite like clockwork”, and lineages rose and fell the way lineages do.

Then an agency filled in a form, and a health minister read the form number out on television, and the word variant — which until that afternoon had meant a base that differs from the reference — became the word for a lineage somebody was frightened of. Nothing underneath it changed. There was now a category with an entry criterion of roughly we are worried, sitting alongside two words that already existed. “There is huge confusion over a lineage, a variant, a strain”, and none of the three is defined tightly enough to hand to a computer, and all three get used interchangeably by people who ought to know better, which included us.

The pressure came from the wrong direction, too. A name had to exist within hours of a notice going out, because the press cannot report a thing it cannot say, and a phylogenetic position is not sayable. Country names were the obvious fix and the one everybody had been warned off: “when naming these kind of things, you want to get away from country names because you don’t want to stigmatise things”, Spanish flu being the standing example of a label that outlives everything true about it. The South African group that described the second lineage of concern were careful not to have their country’s name attached to it. It was called the South African variant anyway, for about a year, by everyone.

The second week of January produced a worse version of the same problem. A Japanese group had sequenced travellers arriving from Brazil, found a set of changes that looked serious, and made it public. For a day or two the entire evidential basis available to anybody outside that group was “a screenshot of a tweet”, in Japanese, put through Google Translate — “luckily someone in my group speaks Japanese” — with nothing deposited anywhere a reader could go and check, because depositing takes time.

Then the Brazilian groups published, and there turned out to be two lineages in play, later P.1 and P.2: one carrying the whole set of changes people were frightened of, one carrying a single change out of that set. Both were the Brazilian variant for a few days. The press picked up the wrong one and ran a scare about something that did not exist, and it was the naming that did that rather than the virus. Underneath the noise, the mutations themselves were being checked by hand, because “you have to triple check this stuff before you go and tell clinicians there’s a problem”.

The sharpest question anybody asked that month came from Orla Geraghty*, a research librarian at an English university, who had come on to have the confusion over new variants cleared up, on the assumption that metadata of this sort was well established in the field.

“How much more confusion could you create by labelling the exact same thing so many different ways?”

The honest answer was that there were no rules yet for this particular virus, that everything was being done in hours rather than years, and that “even people in the field don’t understand most of this”. It was not a good answer. It was the true one.

A name you can catch

A name can say where a thing sits in the family tree, or it can point at something it carries, and the trouble starts when the two stop lining up.

B.1.1.7 is a position in a tree. It says who the ancestors were and nothing else. 20I/501Y.V1 wears one of its mutations in its name — the 20 is the year, the letter is borrowed off the hurricane naming scheme, and 501Y is the substitution people thought was driving transmission. Both schemes were describing the same object in January 2021, and the difference mattered because “we’re seeing similar changes in similar parts of the genomes arising totally independently in different parts of the world”. Group by mutation and you’ll merge unrelated lineages that happened to arrive at the same answer. Go by lineage name alone and you won’t see that the same answer arrived more than once. It takes the tree and the mutations read together to see both.

Enteric bacteriology got there long before and made a worse job of it, because the clinical names were all coined before anybody knew the organisms were one organism. There were microbes causing meningitis in neonates, microbes in the urinary tract and microbes causing gut disease, all described separately, and it was Escherich who suggested these might be the same thing. So enterotoxigenic, enteropathogenic, enterohaemorrhagic, uropathogenic and Shiga-toxigenic E. coli — ETEC, EPEC, EHEC, UPEC and STEC — are not groups. They’re syndromes with a bacterium attached, and everything since has been an attempt to square them with the genealogy, which is the sense in which “we’ve sort of gone ass backwards”.

Which is fine right up until you have a genome on your screen and somebody asks what it is. Shiga toxin is carried on a prophage, so the answer to that one is trivial — “Shiga toxigenic, you find shiga toxin, you’re done” — and it tells you nothing about ancestry, because the phage moves and the name moves with it. EPEC takes a bit more work: look for the locus of enterocyte effacement, which encodes the type three secretion you need to make an attaching and effacing lesion, and you’re probably looking at one. EHEC against STEC is trickier, because EHEC is more about the clinical presentation than about any particular genomic feature. Sequence typing doesn’t rescue it — for the clonal things, O157 or ST131, a sequence type (ST) is fixed and useful; for a random E. coli your ST is “not going to be necessarily a good predictor of its pathology”. Germany in 2011 made the point: an enteroaggregative lineage picked up Shiga toxin and became something else, enteroaggregative and Shiga-toxigenic at once.

Salmonella looks like the well-behaved cousin and is not, and the reason is the argument about species boundaries running one level down: a serovar is a label somebody drew before anybody could check it, and the checking has not been kind. There are over 2,500 serovars, which are antigenic types rather than clades, and the twenty that show up in clinic really are tight, static, boringly consistent buckets. Step outside them, to the ones that live in rivers and reptiles, and the whole thing loosens up and starts shuffling genomic material about like an E. coli. The neatness of the clinical serovars is partly real and partly an artefact of what we sequence, and in 2021 what we sequenced leaned heavily towards Typhimurium, Enteritidis and Typhi. The other part may be that the intermediates that would show how the serovars arose are no longer there to find: “All of the missing links are gone, they’re dead.” Either the chain of events that produced them can’t be reconstructed any more, or it’ll take a lot more sequencing to find out.

Kauffmann and White ran out of diseases

The Salmonella scheme — Kauffmann and White’s, built up from the 1930s and still carrying their names — is the best worked example in the field of what happens when a naming system meets its own scale, and it is “the classic software engineering problem of making it up as you’re going along”.

Version one was descriptive: name the serovar after the disease. Typhimurium is typhoid in mice. Choleraesuis is cholera in pigs — “it causes cholera in pigs, it kind of doesn’t really, but there you go” — and Abortusequi is “another one that explains what it does, what kind of disease it presents, causes abortion in horses”. Then the diseases ran out, which they do, because there are far more serovars than there are syndromes. So version two was geographic: name it after wherever the isolate came from. Dublin, Kentucky, Brisbane, Mississippi. You can amuse yourself finding out whether your home town is a serovar, and most people’s is. The place is only where somebody happened to swab, it “doesn’t mean these originated from there”, and the names are not fixed either — a serovar was renamed in a 2026 erratum, because a committee decided the old name broke a rule. Which is the same problem the virus people were having, eighty years earlier, with the same fix and the same failure.

What the scheme mostly stayed away from was naming anything after a person. Mosquitoes and Plasmodium are full of species carrying the surname of whoever worked on them a century ago, and the question that raises — “do you really want to be associated with this deadly pathogen or vector” — answers itself for most people. The reason given on the Salmonella side is procedural rather than tactful: “that kind of naming is for species”, and a serovar is not one.

Version three is the interesting one. Outside subspecies enterica the antigenic formulas stop meaning anything useful — too many combinations, and the serovars sharing one turn out to have nothing else in common — so the naming committee stopped issuing names out there and told everybody to use the formula instead. That is a string of letters, colons and numbers that describes what the organism reacts with and nothing about the organism, which is “as bad as a computer scientist saying yeah please use the md5 hash” as the identifier, and the reply to that was: basically yes.

What makes it a properly bureaucratic move is that the unusability was the point. The reasoning ran “stop using this we’re going to make it more difficult because we don’t want you to use” it — the serology is not helping you, so we will make the alternative annoying enough that you give up the concept. It did not work. It never works, for the reason the whole chapter keeps arriving at: “we all like a good name”, and a name you can say to a colleague beats a correct string every single time.

Meanwhile the names that survived come with their own compliance burden, because a serovar is not a taxonomic rank and must not be dressed as one. Salmonella enterica subspecies enterica serovar Typhimurium: italics on the genus and species, roman on the word subspecies, italics again, roman on the serovar. Get it wrong in print in front of the wrong person and you will hear about it. There is a white paper from the Centers for Disease Control and Prevention that sets out exactly how to do this, and the definitive statement about how well any of it works is:

“I have a whole document somewhere from people I work with the CDC on how to name salmonella appropriately, and I lost it.”

Somebody has to check the label

Which brings us, eventually, to the hamburgers.

“The UK had the hamburgers that didn’t come from cows” is how Callum Devlin* summarised the horsemeat scandal a decade later. That was a label being checked against what was actually in the packet.

But a label can also fail because the terminology has moved on. Lekha Pillai*, a bioinformatician at a large American culture collection, described strains deposited under names that no longer matched their genomic classifications. A deposit from the 1990s could acquire a contamination flag because its old name and its current classification disagreed. The flag didn’t establish that the culture was contaminated.

Lekha’s group then had to check whether the name had changed before deciding what the mismatch meant. The material might be fine; the label needed its history attached. That’s a practical reason to keep track of old names, rather than merely replacing them.

Where it stands

By 2023 the quarrel about names had stopped being disciplinary and become an international one, with two live models to argue over. The World Health Organisation’s guiding principles for data sharing carried one option that was, in the room, unmistakable: “it was very clear one was very aimed at GISAID, it was like the GISAID model, but they didn’t say it” — GISAID being the influenza-data initiative that had become the pandemic’s biggest sequence repository, and which grants access on terms rather than freely. The other was the sequence database model, share everything with no restrictions. Underneath sits a genuine collision — the equity and data-ownership concerns that made GISAID exist are real, and they run against a genomics culture whose entire momentum has been, in Callum’s words, “entirely open data, open tools, open resources without any restrictions”.

Callum’s experience of sharing data ran against what we’d have expected. Agricultural sequence data was often harder to get hold of than human clinical data.

“There’s a lot more laws for the human side, but there’s a lot more lawyers in the agri-food business.”

Callum also objected to how loosely One Health was being used. It was a term “which has become a critiqued word, if not a meaningless buzzword in some contexts”, and he wanted the discussion to reach further than it did, into plant pathogens for a start. A name can fail by being unsayable, and it can fail by being said too often.

What Callum valued in working groups was the boring, specific, useful work rather than another global strategy, and his example was a pipeline with more than one antimicrobial-resistance detection tool in it. We added agreeing what metadata to collect, describing it in a spreadsheet somebody could understand, and look-up tables between the vocabularies already in use. That gives you something to work with, not just a document to read and file.

Which raises, unavoidably by 2023, whether a language model could just do the ontology instead. Put to Nia Llewellyn*, one of the organisers, twice, with a microphone:

“Gentlemen, I am dropping this mic, taking no more questions at this time.”

The virus naming did get resolved, and it resolved in the direction the whole chapter predicts. In January 2021 the hope was that the WHO would bring people together on it. On the 31st of May that year it did, by assigning Greek letters: B.1.1.7 became Alpha, B.1.351 became Beta, B.1.617.2 became Delta. Nu and Xi were skipped, one for sounding like new and one for being a common surname, which is the same stigma calculation that pushed everyone off country names in the first place. Alpha to Omicron is fifteen letters and thirteen labels, the two skips visible in the arithmetic. The letters were public-facing labels laid over the lineage names, never a replacement for them, and Omicron is where they stopped. That wasn’t the alphabet running out. In March 2023 the WHO began tracking Omicron’s sublineages individually and kept Greek letters for any new variant of concern, and a new lineage didn’t automatically qualify. Everything since has been a Pango lineage with a number on it, and nobody outside the field says one out loud with any confidence.

So the sayable name and the true name went their separate ways, which is exactly what Salmonella did eighty years earlier and for exactly the same reason. It hasn’t stopped anybody working. The one prediction from January 2021 that came in dead on was Orla’s throwaway one, that the media would eventually burn out on new variants and stop caring. They did.

There is a Salmonella serovar called Weitmar. Until a committee looked at it, it was called Scharmann1. Nothing about the bacterium moved.

Notes

  1. Téllez-Castillo CJ, Rekendt AK, Kollberg-Dix S, Pra-Mio L, Scharmann C. (2026) Erratum and republication: identification of a new Salmonella serovar — Salmonella Weitmar (8:z41:1,5). GMS Infectious Diseases 14. doi:10.3205/id000103. The WHO Collaborating Centre for Reference and Research on Salmonella, Institut Pasteur, ruled that Scharmann was a personal name and so did not comply with current nomenclature rules.

Pangenomes

The first thing anyone noticed in the bug report was not the bug. It was that GNU parallel had politely asked to be cited, and had been ignored, sixty-two consecutive times. “You’ve run parallel 60 times, you’ve run parallel 61 times, you’ve run parallel 62 times”. Somewhere under that stack of increasingly plaintive notices was the actual error, which read “sequence without letters could not guess alphabet”, and which took about four seconds to guess: an input with no sequence in it, quite possibly a GFF file downloaded from the National Center for Biotechnology Information (NCBI) in the flavour that leaves the sequence off the bottom. An easy mistake, a common one, and nothing whatever to do with GNU parallel.

That was one episode: sitting down with the public GitHub issue tracker of a piece of software one of us wrote, reading the reports out loud, and guessing the cause before scrolling down to the answer. It is not a normal thing to do. Nobody does the equivalent with a paper, which is a snapshot in time.

What a pangenome is for, and why it was hard

Sequence a few hundred isolates of the same organism, assemble them, and the assemblies do not agree with each other. Some of the disagreement is plasmids, phage, integrons — things that arrive and leave. If all you want is a tree, you can map everything to a reference and ignore the rest, but the rest is frequently the interesting part, and to get at it you need to know which genes are in everything and which are in some things. The genes in everything are the ones that define the organism at all, “what makes a Salmonella an actual Salmonella”, give or take a couple of thousand genes. The rest is what it is carrying this week.

Around 2013 that was a genuinely difficult calculation to do at scale. The Sanger Institute had sample collections far larger than anything the available tools could chew through, and the tools themselves were unfriendly — OrthoMCL had by then acquired a MySQL database, which turned a simple question into an administrative project. The literature at the time topped out at pangenomes of eighty or a hundred bacteria. The target was thousands to tens of thousands.

What got it there was one unglamorous trick. The method underneath is an all-against-all comparison of every gene against every other gene, which is quadratic and therefore hopeless. Cluster the genes first with CD-HIT and the comparison shrinks enormously, because a species only has so many distinct genes in it — add more genomes and you are mostly adding copies of things you have already seen. The curve stops being n by n and starts looking more like a line, and the whole thing becomes possible on one machine over a weekend.

The weekend on one machine was not the plan. Roary is a collection of scripts that call each other through job runners, and the job runners are there so the work can be spread across a cluster — thousands of jobs at once, which is what an institute with an LSF queue in 2013 had to hand, and which took a long time to get right. What everybody did instead was find a box with sixty-four cores, ask for sixty-four threads, and never go anywhere near “the awesome power of HPC”. So the feature stayed in and stopped being mentioned: “that whole functionality is still there, but it’s kind of hidden away and no longer advertised”. Seven years later one of us heard about it for the first time on a recording — “wait I had no idea about this” — and got the honest answer, which is that it still works, it still speaks only LSF, and “you can actually use it but you don’t really need to”.

Internally it was called ‘The Pangenome Pipeline’, and the scripts were called things like create_pangenome. That is what it was still called when other people started wanting it, all happening at a time when Andrew, the developer, was having cancer treatment.

Core, accessory, and the line between them

The split into core and accessory sounds like a property of the bacteria. It is a threshold somebody chose, and everything awkward about it follows from that.

Start with the obvious definition: a core gene is one present in every genome in the set, and an accessory gene is one present in some of them and not others. The problem is that assemblers do not assemble everything. A contig break falls in the wrong place, a gene that is unquestionably there fails to be called, and one bad assembly out of five hundred removes a gene from the core that belongs in it. Demand a hundred per cent and what you have measured is your worst assembly.

“So you have to allow a little bit of fuzziness.”

So the core is defined at ninety-nine per cent, and now you have arithmetic. Whole genomes come in whole numbers, so the count of genomes a gene is allowed to be missing from ticks up by one every hundred genomes you add, and the core genome size jumps at each tick. This turned up in the FAQ as a real question from a real user — “why is there a sudden increase in core genome size every hundred genomes” — asked by somebody who had drawn the curve, seen a step in it, and quite reasonably started thinking about biology. It is rounding.

Then there is the same gene twice. Bacteria carry multiple copies of things, and a naive clustering merges them into one entry, which is wrong. What Roary does instead is look at the context, using the five genes upstream and the five downstream as a fingerprint, so two copies sitting in different places in the genome come out as two separate entries. People complain about this, because it means the same gene name appears more than once in their core. The complaint has the biology backwards. A copy inserted somewhere else has a different history from the original, and “they are going to be different, if it’s inserted into a different location they are separate things”. You can switch the splitting off with a flag. Mostly you should not.

The third thing that moves the line is annotation, and it is the one people fall into hardest. Roary takes Prokka output and, deliberately, nothing else, because generic annotation files are a royal pain and because the alternative is worse. If you download fifty genomes from GenBank that were annotated by fifty different pipelines, the gene callers disagree about where genes start, the clustering sees those disagreements as real, and your pangenome is a picture of annotation software. Annotate everything yourself, the same way, in one go, and then compare. There is a real loss buried in that advice, because “it’s more important that it’s consistent rather than having sort of nuance or accuracy involved”, and a lot of painstaking curated annotation gets scrubbed off and replaced with the automatic kind. “There’s a lot of secret information that just disappears along the way”.

Which brings you to the report that is the best argument for understanding any of this. A user with twenty-five genomes of closely related bacteria, all annotated, got back zero core genes, zero soft core genes — the band just below the core, present in most but not all — and three thousand eight hundred genes in clouds of their own, meaning found in only one or two. Every isolate almost entirely unique. There are two numbers worth carrying around here: a gram-positive and a gram-negative thrown into the same analysis still share a few hundred genes, and a random E. coli against a random Salmonella shares around two thousand. Zero is not a biological result. Our guess was that the assemblies were garbage — probably all of them, contigs so short that genes were never called reliably — and the first thing we’d have asked for is QUAST run over the lot.

Read the other way, that makes the core genome size a QC instrument, and a cheap one. Contamination shows up in it immediately: a hundred and fifty core genes across a set that ought to be one species means somebody has accidentally built the pangenome of every gram-negative organism they had lying about. It gets used as a first check on incoming data for exactly that reason. The FAQ carries the same point in one line — “I haven’t done any QC on my sequencing data and the pangenome looks very strange”, answered with garbage in, garbage out — and that is the most common support request there is.

The name

It could not ship as the pangenome pipeline. The name it got was Roary1, after Roary the Racing Car, which is a children’s programme, and also after a son with the same name spelled slightly differently — later introduced on air, deadpan, as “this is the original Roary”. Racing cars go fast. Nobody was claiming the name did any work beyond that.

Except that it did, and the reason is the least romantic possible one: “It was a unique namespace.” A word nobody else had used meant that a user with a problem could type the name of the tool and a description of their error into a search engine and land on an answer. That is not branding, it is a support strategy, and it is the difference between a question that gets answered by the internet and a question that gets answered by you, personally, on a Tuesday.

Where it stands

The Roary FAQ was written the same way, one entry at a time, every time an unusual support request came in. It reads as comedy — “It’s a barrel of laughs” — and it is, partly. There is an entry for the person who tried to run the Perl scripts with Python. There is an entry answering whether the author will analyse somebody’s data for them: “Sure, pay me a lot of money and I’ll do it.” But it is also the only honest documentation of what maintaining a widely used tool consists of, which is the same handful of questions forever, most of them not about the tool.

The people asking are usually not bioinformaticians. They have been handed a set of data because somebody senior wants a tree for a paper, and they would rather not learn the field to get one. That is treated with less sympathy than you might expect, and the analogy offered was physical: “I can’t walk into a pathogen lab and start sloshing around chemicals without knowing what I’m doing. I will be shot by health and safety.”

The hardest ones to answer are not a single error but two or three hundred lines of log. Somebody installed through cpanm, which fetches the Perl dependencies and none of the system libraries those dependencies are built against, and the report is the whole log: modules that installed, modules that did not, “dependencies of dependencies of dependencies which are failing”, and nothing anywhere saying which of them went wrong first. Working down it takes three of us a couple of minutes. BioPerl failed, so BioPerl was never installed — except that modules with nothing to do with BioPerl failed too, which puts the problem lower than that: no zlib, or no compiler, because “on a lot of systems they don’t install compilers by default”. Then a database module scrolls past and the guess narrows again. It is a decent diagnosis and it is close to useless from a browser tab, because with errors like these “you actually have to be physically at the machine to try and debug it”.

The other half of the sympathy shortage points inward. A stack trace that sends the user hunting through source code is not much help. We suggested checking dependencies before starting, but Roary already did that. A dependency could be present and still crash when called. That is harder to catch than a missing executable.

A useful bug report still needs enough detail for somebody else to reproduce the failure: how you installed the program, what you gave it and what happened, because “other people aren’t mind readers”. GitHub issue templates let us ask those questions on the form rather than hoping.

And then there is the one that cannot be fixed. Somebody had run three thousand isolates and wanted to add one more. There is no way to do that. The clustering needs everything at once, so adding a single genome means running the whole thing again from the beginning, and the fix has been on the list since the start: “it is actually something that I’ve known from the start I want to put in but it’s a huge amount of engineering work”. Newer pangenome tools do handle it, and the reason they do is not that they are cleverer.

“I know other people have gone and done it — clean implementations of pangenome stuff — and it’s worked because they’ve thought about it from the beginning but I didn’t.”

That is a design decision made in 2013, being explained without excuses by the person who made it, to an audience of people who use the software. It is rarer than it should be.

One judgement from the same session has aged. Asked in late 2022 for an Apple Silicon build, the answer was that the machines were “still an obscure platform that only rich people can afford”. Within months Apple had finished moving its whole line across, so the platform was not obscure and buying into it was not a choice, and the objection quietly became the ordinary one of not owning the hardware you are being asked to support.

“I feel like this is our version of celebrities read mean tweets or something.”

Except that none of them are mean. They are twenty-five bad assemblies, an absent compiler, a GFF with nothing under it, and one person who wanted to add a single genome to three thousand. The answer to that last one was to start again from the beginning, and nobody enjoyed giving it.

Notes

  1. Page AJ, Cummins CA, Hunt M, et al. (2015) Roary: rapid large-scale prokaryote pan genome analysis. Bioinformatics 31(22):3691-3. doi:10.1093/bioinformatics/btv421

Mobile Genetic Elements

A Vibrio came in that needed typing, and it had “three phages back to back to back in it”.

Short reads could not get across them. Nothing could get across them, because the three phages were near enough identical and a repeat you cannot span is a repeat you cannot place, so the assembly came apart at exactly the point the assembly was about. PacBio was only just becoming available. The strain was finally spanned about a year into the study, and the whole account of that year is four words long.

“It was so hard.”

The definition does not help you

A mobile genetic element is a piece of DNA that can move itself between genomes, or between chromosomes, and the taxonomy of them is a list that keeps going: bacteriophages that integrate as prophages and later cut themselves back out, transposons, plasmids, genomic islands, integrative conjugative elements, and integrons, which are capture platforms that collect resistance genes into a row and add to it. Things that look like plasmids and then dive into the chromosome. And then, because none of these categories are enforced by anything, remixes — “you can have something that’s sort of phage looking, but isn’t”, and prophages missing half their own excision machinery that borrow the rest from a neighbour when they want to leave.

The working definition is therefore circular — “As long as there’s some sort of mechanism for it to transfer, then it’s a mobile genetic element” — and whether there is such a mechanism is usually the thing you were trying to establish. There is no signature to grep for. There is a family resemblance and a lot of exceptions.

The standard move is to run PlasmidFinder and read off the incompatibility type, which is a real answer as far as it goes. It goes about as far as the replicon — the stretch that carries the machinery for copying itself, and the part a typing scheme can actually see. Those same replicons turn up sitting in chromosomes as well as on free plasmids, and from short reads there is no way to separate the two — “you have no idea whether this is something that’s accidentally been misassembled” into the chromosome by your assembler, or whether a plasmid genuinely integrated there and stayed. The two look identical on the screen, and the answer to which one you have is “Difficult to say”.

It is also worth knowing before you start that most of what you find will be doing nothing at all. Take two random bacteria and the differences between them will include a phage that confers no advantage on anybody and is simply along for the ride, and plasmids in the same condition are common enough to have their own name — “they call them cryptic plasmids like they don’t know why they’re there”, which is an admirably direct piece of terminology for a field that could have called them something flattering instead.

The honest description of how the field handles all this arrived early and has not improved.

“We kind of pretend, a lot of us pretend, that it’s not actually there and try not to think of it too much and pretend everything is just nice and tidy sequence types.”

That pretence is not laziness. It is what nearly every method in this book is built on. A sequence type, a core genome, a tree — all of them assume a genome has one history, inherited downwards, and mobile elements are the standing evidence that a genome is a bundle of histories that happen to be travelling together this week. It is easier to draw the tidy version. The tidy version is also, for most of a genome most of the time, right enough to work with, which is precisely what makes the exceptions so expensive.

When the exception is the whole story

The typhoid work is where this stops being a modelling preference.

Salmonella Typhi kills on the order of a hundred thousand people a year, mostly people least able to reach healthcare. H58 is the lineage behind most multi-drug-resistant typhoid, already the one everybody watches. In samples from outbreaks in Pakistan, it had acquired a plasmid carrying additional resistance genes1. Farid Tariq* set out what made this extensively drug-resistant typhoid: resistance to the older first-line drugs, plus fluoroquinolones and cephalosporins, leaving azithromycin and carbapenems as the treatment options.

The detail that matters most is the least dramatic one. Plasmids move in and out, and something that moved in can be kicked back out again when the cost of carrying it stops being worth paying. But some of the resistance genes had integrated into the chromosome, and “when they then get integrated, they’re very difficult to then get rid of”. The mobile element does not have to stay mobile. It can arrive, unpack, and become part of the thing it arrived in.

Working backwards through the public archives, Andrew traced the plasmid to an E. coli from cattle, caught by international surveillance. The organism that spread in people in Pakistan and the DNA that made it extensively drug resistant have different ancestries, different hosts and different stories, and no phylogeny of Salmonella was ever going to show you the second one.

Every step of the pipeline is a filter against them

Here is the mechanical problem, and it is worth being precise about, because the failure is silent.

Start with what identification actually needs, which is not the element but the ground around it. A region gets called a prophage or an integrative element on the strength of what sits either side of it and in what order, because “just finding a singular transposon in the vacuum of gene space doesn’t tell you anything”. So you need an assembly, and the assembly has to be good enough to hold the element and its neighbours on the same contig.

Mobile elements are repetitive. Prophages come in multiples — some E. coli carry ten or twenty prophage-looking regions — and insertion sequences come in far worse multiples than that. Pertussis has an IS element you can find “hundreds or maybe thousands of times in the genome”. Shigella is full of them, and the practical consequence is that “usually your N50s are pretty lousy”. Every copy of a repeat longer than your read is a place the assembler cannot decide between two continuations, so “each of those will probably be a break point”. The mobile fraction of the genome is, almost by definition, the fraction your assembly is worst at.

Two tricks find them anyway, and both work by looking at something other than the sequence. Map your reads back onto your own assembly and plot the coverage: where a repeat has been collapsed into one copy, “you’ll have a massive spike over a certain region of the genome”, and you do not need to know what any of those genes do to know something is there more than once. And look at the GC content, because prophages tend to run AT-rich against the background chromosome, so a foreign chunk often shows up as a step in a plot before it shows up in an annotation.

Long reads were supposed to end all of this, and the first thing they did was make it worse. The early algorithms — some of which are still in use — took the shorter of the long reads and used them to correct the longer ones, on a length threshold, so that “anything below 3000 bases would be used to correct all the other long ones”. You got beautiful high-quality long reads. You also got a rule that quietly removed anything whose reads all fell below the threshold, and some mobile genetic elements are very short.

“You’d have disappearing plasmids.”

Nothing errors. Nothing warns you. The plasmid is not in the output because its reads were used up correcting the longer ones, and the only way to notice is to have expected it. This is the case that should make anyone nervous about a pipeline they did not read: “bioinformatics doesn’t solve everything and it can introduce extra errors”, and the errors it introduces are subtractions, which are the hardest kind to see.

The ones we fire on purpose

The compensation for all of this is that mobile elements are also the best tools in the building, because a thing that inserts itself into DNA at random is exactly what you want if your question is which genes a bacterium cannot live without.

Holly Penrose* starts where everybody started. A transposon is a length of DNA that jumps — the “blue and yellow corn” everyone met at school, before anybody mentioned that the mechanism would one day be a reagent you could buy. Insert them as randomly as you can into a population of bacteria, so that each cell has a different gene interrupted, grow the population under whatever condition you care about, then sequence the survivors and look at where the insertions landed. Anywhere you find an insertion is a gene the organism could do without. The genes with no insertions in them are the essential ones, because “if the genes are essential, they have to be there, there can’t be a knockout” — the mutants that lost them are not in the culture to be sequenced. The whole method is an absence, read carefully, and comes with many names like TraDIS, Tn-Seq and INSeq.

Bacterial immunity got repurposed the same way. CRISPR arrays are a record of foreign DNA the lineage has previously run into. The genes that go with them are well conserved, which makes the arrays easy to find, and the spacers and repeats are short enough to assemble cleanly, so people typed on them. Spoligotyping in tuberculosis is exactly that: a standard typing method built entirely out of a bacterium’s filing system for old phage encounters. Whether it tracks the organism’s own history depends on the organism — in some Salmonella serovars the spacers have followed the lineages for a long time — but CRISPR typing was never as reliable for surveillance as SNPs or multilocus sequence typing, and it went the way of most pre-genomic typing the moment a whole genome became the easier thing to produce.

And phages themselves can be turned on bacteria. Lucy Farrant* had done some phage work on exactly that, using phages that attack bacteria to clear an infection. The theory is that something which has spent its whole existence killing a particular organism might be good at killing it. The catch is the filing system from a moment ago: a bacterium that has recorded a phage is resistant to it, although, as Holly pointed out, not all bacteria have the system.

Anyway. The point is that the elements that ruin an assembly and the elements fired deliberately into a genome a million times are the same elements, and it is only the direction of travel that differs.

Where it stands

Some of this has aged well. Long reads did deliver on the repeats this chapter opened with, although a small plasmid is still something to go and look for rather than assume, and adaptive sampling now lets a run decide what to keep while it is still running. Sequencers went to a lot of places during the pandemic that did not have them before — a MinION and a laptop posted to a national reference lab in Zimbabwe, a first long-read run done for COVID, and then the same equipment and the same training available for Salmonella, which was what the project had been for in the first place. That is probably the most durable thing the period produced.

One prediction from April 2023 has not landed the way it was expected to. The next big step was going to be culture-free diagnostics: a sample in, no growth step, “you get a result in a few minutes” telling you which antibiotics will work. The nearest thing to it is in tuberculosis. The World Health Organization’s 2024 diagnostic guidelines conditionally recommend targeted sequencing straight from respiratory samples to detect drug resistance in people already confirmed to have pulmonary TB, which takes the weeks of culture out of the resistance result, although culture has not gone away and the answer takes days rather than minutes. The rest of it — a primary sample, no culture, an antibiogram while the patient is still in the department — exists in a small number of places and is not what happens when you go to hospital.

And there is a caution attached to all of it that has aged very well indeed, offered to a room of PhD students by one of us, who did a machine learning PhD long before that was a fashionable thing to have. The data is not good enough yet to mine, the enthusiasm is running ahead of it, and the distinguishing feature of a weak signal in this field is not that nobody finds it.

“Give a PI a random set of five genes, and they’ll make up a story saying exactly why that has happened.”

The duff

You go and look. That is the actual answer, and everyone on every one of these conversations converges on it: assemble, annotate roughly, then open the thing in a genome browser and read it, because the tools get you most of the way and the last stretch is a person deciding what they are looking at. “I trust no algorithms at all” is a strong statement of a position nobody really argues with. It is an easier position to hold when everyone in the room agrees that “mobile genetic elements are definitely a wild west still” — run three phage finders over the same genome, get three answers, then work out which one you believe.

And sometimes you look, and there is a phage tail, and a helix-turn-helix, and a run of genes between them, and you put the sequence into nr and get back nothing at all — “this has never been seen, no one has ever deposited this ever”. You try more sensitive alignments. You try profile databases, domain scans, structure prediction. Still nothing. At which point the options reduce to going and talking to someone with a wet lab, since “you have to do actual work at that point”, or writing it off.

“This is just an unknown, a duff.”

The isolate still goes on a branch. The branch still gets drawn.

Notes

  1. Klemm EJ, Shakoor S, Page AJ, et al. (2018) Emergence of an extensively drug-resistant Salmonella enterica serovar Typhi clone harboring a promiscuous plasmid encoding resistance to fluoroquinolones and third-generation cephalosporins. mBio 9(1):e00105-18. doi:10.1128/mBio.00105-18

Resistance

The plasmid is 4.4 kilobases and carries three genes. One replicates it, one helps it move, and the third is an efflux pump that pushes tetracycline back out of the cell before the drug can do anything. It has spread across more than thirty clonal complexes of Staphylococcus aureus — at least thirty separate horizontal transfers, since the plasmid is younger than the lineages carrying it — and it assembles out of short reads as a single contig, so a BLAST search either hits it over nearly its whole length or does not.

Isolates that carry it are not reliably resistant to tetracycline.

“We see some isolates that have the plasmid, which are still totally susceptible to the drug, and some which are very highly resistant.”

Ffion Howells* thinks copy number is at least part of the explanation. More copies of the plasmid, more copies of the pump, more drug leaving the cell before it binds anything. In both cases the gene call is the same, and the call is the thing that gets reported. Present. Yes.

A gene call is a yes or a no. A drug is a dose. Nearly everything difficult here lives in the gap between those two sentences, and the clinic is on the far side of it.

What the table says

The ordinary workflow is short. Assemble, run ResFinder and PointFinder over the contigs, get back a list of gene names, and for a good number of determinants that really is the whole answer — sulfonamide is about as easy a call as exists, and if an isolate carries mphA and ermB it will be resistant to azithromycin at a high level. Nobody is claiming the table is useless. The claim is narrower and worse: the table is a different kind of statement from the one a doctor is waiting for.

Take beta-lactam resistance, which Beth Stannard* unpicked for us in 2020, and which is not one mechanism but a stack of them. Several of the Enterobacteriaceae carry AmpC on the chromosome — E. coli does, Klebsiella and Salmonella largely don’t — and in E. coli small differences in the promoter shift how much enzyme gets made and therefore how resistant the cell is. On top of that sit the acquired genes that arrive horizontally, blaOXA and blaTEM and the rest. Behind both, efflux. It is a phenotype assembled out of several parts at once — “if it were a human trait we’d probably call it polygenic” — and a table that reports the parts one line at a time is not reporting the trait. Her verdict on doing it anyway is that “gene presence equals phenotype presence is just kind of artificial.”

Underneath the obvious determinants there is a whole register the tools were never built to see. Regulation. Expression. Small shifts in the minimum inhibitory concentration that come from bacterial defences nobody has yet decided to classify as resistance at all. “There’s a lot of stuff happening underneath to do with regulation”, and none of it produces a hit in a database of gene sequences.

And then the answer gets rounded. Resistant or susceptible, one column, one word, or even just one character — which is convenient for a spreadsheet and is not what the organism is doing.

“This is actually continuous data which we forced to be discrete for the convenience of analysis.”

Asked whether somebody could not simply write the one tool that solved this for everybody, Beth’s reply was that “if you can, you’d never have to work again.” Five years on nobody has had to stop working.

The things a gene call cannot see

Start with what is in the tube. A sample is a population, and the pipeline usually assumes it is an organism. Work on two-hundred-year-old tuberculosis genomes from Hungarian mummies stalled for a while on results that made no clinical sense, until it became clear that a single sample held two genotypes at roughly fifty-fifty. The pipeline had already thrown that away: it was calling single-nucleotide polymorphisms (SNPs) and discarding anything supported below seventy per cent, so a signal sitting at half the reads did not exist as far as the software was concerned. The threshold was doing exactly what it had been written to do.

The clinical version of that arithmetic, as Graham Thorne* gives it, is unforgiving. Suppose the report is right and ninety-five per cent of what is there is susceptible. Give the drug that the other five per cent resists and the resistant five per cent will probably take over the culture. Minority populations can decide the outcome, and a gene table has no column for them.

There is a nice moment in the middle of that argument where an assumption gets overturned live. Graham’s clinical understanding was that you get tuberculosis once, and his first reaction to the mixture was “you don’t catch TB twice”. Then, out loud, within the same answer: “contrary to what I said earlier people do get infected twice with TB” — because in a setting where a quarter of patients show evidence of mixed infection, and where exposure happens daily in crowded rooms, why on earth would you not.

Populations also move in time, inside one person. Paige Whitlock’s* sequential isolates from patients with recurrent methicillin-resistant S. aureus bloodstream infections come back a handful of SNPs apart, close enough to point to one source population rather than a string of fresh infections from around the hospital. Whether that source is the patient or somebody close to them is a separate question, and Paige leans towards the patient. What tips her is where the differences sit: in a resistance-associated gene, the same change turning up again in the later isolate. The argument is not statistical, it is a question about who was on the drug — “if your housemate isn’t getting treated with antibiotics”, it becomes fair to ask whether the selection for that particular mutation happened anywhere else. Reinfection from a contact isn’t ruled out. But the likelier story, on her account, is that something was not cleared, it was treated, and it came back having moved: resistance being made on one person, with no transfer event to detect and no new gene to find.

Which puts weight on a decision made at the bench long before any of this: how many colonies to pick. Two isolates taken from the same patient can differ substantially, and the variation turns up even between preps of what is nominally one clone. Then “for financial reasons we’re only going to do one isolate per patient”, and the population you were trying to characterise is one colony that grew well on the day.

The other half of what the call cannot see is where the gene is sitting — the mobile element problem arriving from the clinical side, with a patient on the end of it. Shigella comes off a standard short-read library in several hundred pieces — insertion sequences again, doing to the graph what they do everywhere else in this book — and the breaks fall precisely in the places resistance likes to live. You recover the integron and lose what it was on. In 2020 the answer to this was long reads, and long reads have in fact arrived for isolates. Metagenomes are harder, and Patrick Garvey’s* first question is whether an antimicrobial-resistance gene survives assembly at all — whether the “AMR gene makes it into the contigs, which a lot of times it won’t.” His own attempt goes at the assembly graph directly. Querying it for the gene already recovers more genes, with some of their context, than searching contigs does, while still recalling fewer than reads alone. The next step, still a plan when we spoke, is to read the structure around each hit — the flanking segments, GC shifts, the likely taxonomy of nearby pieces — as evidence of a transfer. Graph work also attracts machine learning, and his honest description of that, from someone with one foot in computer science and the other in microbiology, is that “One wants to shove all data known to man through machine learning and the other wants nothing to do with machine learning and thinks all of it is Tosh.”

All of it then gets checked against a phenotype, and the phenotype in question is a minimum inhibitory concentration (MIC): a number off a plate. What bothers Beth about the benchmarking studies is that they ask whether the genomic prediction recapitulates the MIC. Nobody on a ward is treating a plate.

“The gold standard should really be treatment response and treatment failure rather than whether or not the bugs behave the same in the lab.”

Building that cohort means following people over years, recording what they were given and what happened, with enough of them to see past everything else that decides whether somebody recovers. It would answer the question. In 2020 none of us knew how to build it.

What they were writing in 1948

The first modern tetracycline, aureomycin, was described in a 1948 paper1 that Ffion called her favourite, and that is a considerably better read than anything published this year. Old antimicrobial literature was written by people who were plainly enjoying themselves: “if you want to describe a bacteria as an ultra-mold, I guess you could do that in 1948”, and a species you considered underappreciated could be called “the hound dog” without anybody objecting. They also felt free to name a subject in the introduction, say plainly that they would not be discussing it, and then not discuss it — which no journal would now permit, and which is arguably the single greatest loss to scientific writing of the last century.

But the reason to go back is not the prose. What struck Ffion in the early work was how soon the evolutionary framing got waved off. Resistance was written about, from nearly the beginning, as a property of a molecule and a cell:

“A simple interaction, a simple biochemical problem when in fact it’s part of a much larger system of interactions.”

Faisal Siddiqui* has seen what that framing costs now. Tuberculosis line probe assays test for a fixed list of resistance mutations, so a resistant mutation that isn’t on the list doesn’t get treated appropriately — which is exactly the advantage it needed. The result has been outbreaks that amount to “diagnostic-driven selection for mutations that are not in your catalogue”. A gene list is the biochemical framing rendered in TSV, and once it’s acted on it stops merely leaving out population, dose, time and selection. It becomes one of the selection pressures.

Where it stands

Some of this has already been settled by adoption. In Gregory Ashdown’s* account, TB genomics is in clinical use in the UK and was the obvious first case; Salmonella shows good concordance between genotype and phenotype, Klebsiella is decent. Pseudomonas and Acinetobacter show a lot of discordance between the phenotypic result and the genomic one, and have stayed in surveillance rather than diagnostics — which is the right home for a method that is right most of the time.

The obstacles that remain are mostly not the science. Running the pipeline is the easy part. Going from a result to what it means is the hard one, and doing it needs the interpretation written down somewhere that is, in Gregory’s phrase, “not just in someone’s head, in an expert’s head” — the sort of knowledge that different alleles of the same beta-lactamase gene will give you resistance to cephalosporins in one case and carbapenems in another. Nobody has managed to encode that. Meanwhile what reaches a clinician has to be a decision rather than a data dump, and the failure mode there is one Callum Devlin* named: handing over the whole output as though that were generosity, doing our “bioinformatics thing of like, here’s all the data, work it out.”

Accreditation is worse. Callum had been helping a hospital get accreditation for its pathogen genomics, and the agencies he’d dealt with “don’t seem to have any idea what they need to do to accredit use of genomics”, which is a legal and logistical obstacle rather than a technical one and has no bioinformatics fix.

Underneath the agencies sits a smaller version of the same gap, which Cedric Kwok*, then in a hospital laboratory in Qatar, put in terms of licensing. A clinical laboratory runs on licensed staff, and the licence is what gets you onto the floor at all — “if you only have the techniques to do molecular you are not able to get into the operation of the lab”. As far as he knew, nothing covered the sequencing side: he didn’t believe “there’s any licensing for the clinical staff purely based on library preparation for whole-genome sequencing”, in the United States or the United Kingdom. So one constraint on sequencing six thousand cultures a month is cost. The other is staff: conventional microbiology technicians can be hired in numbers, sequencing ones are hard to find, and the protocols, training and licences that would produce them hadn’t been set up.

The most quietly damaging problem is the phenotype data itself. In a hospital lab, the concentrations come off an instrument into a system the sequencing scientist has no access to. What is available to Cedric is paper — “they can print out a hard copy and then give it to me”. Which is why so many public genomes have no MIC attached to them, and why the benchmark that everyone agrees is imperfect is also the benchmark nobody can assemble at scale. A 2022 comparison found massive discordance between prediction methods; the agreement problem is real, and it is hard to fix without the phenotypes that are sitting in PDFs.

Behind all of it is what clinicians will actually do with a result. Prescribing is defensive; the broad-spectrum agent covers everything, so the broad-spectrum agent is what gets given, and a report saying a narrow-spectrum drug would do the job may not change the decision at all. Meanwhile, as Graham points out, the technique that did make it into the clinical microbiology laboratory was matrix-assisted laser desorption/ionisation time-of-flight mass spectrometry — MALDI-TOF — which identifies an organism in minutes for almost nothing and does it by matching a spectrum, not a sequence. “For those of us who like our data to be digital and like a tidy kind of universe” it is an affront, and it won on turnaround and cost while we were arguing about databases.

Nobody knows how much of any of this matters in the aggregate, either. Beth reached for the O’Neill report and its “10 million dead by 2050”, then immediately conceded that nobody really has a good value for it, because the thing being counted is so poorly defined2. What it is, on that account, is a “great global experiment that we’ve been doing with bacterial populations”, and we are measuring the result with instruments built for a different question.

The tools get it wrong in ways that are entirely on us, too. An early study was run on ARDB3, a resistance database everybody knew was uncurated, and it missed an important beta-lactamase; a collaborator caught it before publication. The moral Beth drew afterwards is the flattest one available — “database selection does matter.”

The plasmid is on every continent anyone has looked at — “across all continents, not Antarctica” — and in one common cassette it has been found integrated into the middle of the methicillin-resistance element itself, so a single recombination event delivers both.

A report can say the gene is there. It does not say how many copies, or what else is in the tube, or what the population will look like a week after somebody starts the drug.

Notes

  1. Duggar BM. (1948) Aureomycin: a product of the continuing search for new antibiotics. Annals of the New York Academy of Sciences 51(2):177–81. doi:10.1111/j.1749-6632.1948.tb27262.x
  2. O’Neill J. (2016) Tackling drug-resistant infections globally: final report and recommendations. Review on Antimicrobial Resistance.
  3. Liu B, Pop M. (2009) ARDB — Antibiotic Resistance Genes Database. Nucleic Acids Research 37(Database issue):D443–7. doi:10.1093/nar/gkn656. No longer updated; the study referred to here was corrected before publication and was never itself published.

Bugs With Personalities

The story Kerem Yildirim* tells is of a conference where somebody presented work showing that Campylobacter has a capsule, and the room would not have it. Thirty years of working on the organism and nobody had ever seen one, which an audience member announced on the way to the door.

“I’ve worked on Campylobacter for 30 years and I’ve never seen a capsule. This is horrific. I’m leaving this presentation.”

Then the genome sequence came out with the locus sitting in it — “There’s your 25 to 30 gene loci encoding for the capsule” — and nobody said anything after that.

The reputation you inherit

You take on an organism’s reputation before you have run a single command on it, and the reputations are sturdy. Campylobacter is fussy. Kerem, who has spent a career on it, calls it “a really fussy pedantic bacteria”, and every piece of supporting evidence is about handling rather than biology. Never a model organism. Never a convenient animal model. Genuinely hard to grow — microaerophilic, happiest between five and ten per cent oxygen, and obliged to live in air. Nobody has explained how, and the question has been open long enough to have acquired a name: the great Campylobacter conundrum.

Which is a reputation formed entirely indoors. Outside, the same organism sits in a chicken at ten to the ninth colony-forming units, does the bird no obvious harm, turns up in wild birds and puddles and swimming pools, and needs somewhere between one hundred and five hundred colonies to infect a person. An obstacle-course race in the US produced an outbreak because of the puddles. It also gets into places built to keep it out: Kerem has been through the door at the UK’s leading poultry suppliers, where the security is “like going through a NASA system of security”, and the birds on the other side carry it anyway. Fussy is what it is on a plate.

Mycobacterium tuberculosis holds the opposite reputation from the same kind of source. Ask a geneticist fifteen years ago and tuberculosis — TB from here on — was strictly clonal: one strain in a patient, that same strain passed on, no recombination, no plasmids, nothing much to argue about. Ronan Keogh’s* summary of that view is that “it’s kind of boring for a bacteria”.

And the Salmonella serovars behind invasive disease in rural Gambia are called atypical, which is a reputation built out of an absence. They are the ones that are not Typhimurium and not Enteritidis, and Typhimurium and Enteritidis are the ones the vaccines are aimed at.

None of the three is a fact about an organism. Each is a record of what was measurable when the reputation formed, and all three are in the middle of being embarrassed by better measurement.

Clonality is the clearest case, because the dogma and the diagnostic turn out to be the same object. Culture reports whichever strain grows fastest, so that is the strain that gets called. Penny Ashworth-Clarke* describes a patient who was diagnosed with a drug-sensitive infection, treated, and came back two months later carrying an isoniazid-resistant strain that had taken over: two strains in one lung, and a method that could only ever report the winner of a race it had started itself. Mixed infections in TB are massively under-diagnosed, and under-diagnosed for that reason. Deep sequencing now finds subpopulations that differ between a cavity and the lung, and sputum that does not represent what is underneath it. A good deal of what made TB look boring was a property of the eight-week culture standing between the patient and the sequencer.

What a reference is claiming

The standard TB transmission analysis is to call SNPs against H37Rv, take the distance between two isolates, and call a transmission cluster at fewer than five or twelve SNPs depending on whose paper you are following. Two things in there are quietly load-bearing.

The first is which genome you are actually comparing. About a tenth of it is not in the comparison at all — the repetitive regions, the PE and PPE gene families, discarded before anything is counted — which is why Ronan’s description of the routine is “near whole genome sequencing”. Mutations in what you discarded do not exist, and the mutation rate is not even across the genome anyway, so the rate everyone quotes is a rate over the part that was tractable.

The second is that rate. TB averages about one SNP every three years left to itself, so across the span of a real outbreak the distances collapse, and “you’ll just end up with zero SNPs between two different strains, and then you have no idea if they transmitted to each other”. A cluster in Rwanda has been circulating since the nineties, carries most of the country’s multi-drug-resistant TB, and sits around twelve SNPs wide. The threshold and the clock are the same measurement used twice, and the second use inherits everything that was wrong with the first.

The way out is a model rather than a cutoff — Bayesian transmission methods that take a generation time and an infectivity period as parameters and ask which tree could have produced these genomes. Which requires a generation time. TB’s comes from growth in a lab, roughly a day, and Camille Engel’s* response to that was “that must be a massive fudge”. Lineages five and six grow slower than the others, and nobody is confident they are being fed correctly in the first place.

A reference also decides what can be looked for at all. Hybrid capture baits are designed against H37Rv, which is lineage four. Lineage five strains carry genes that are not in H37Rv, so those genes are not captured, so as far as the assay is concerned they are not there. Lineages five and six are the West African ones.

The other half of the problem is physical. A reference genome drifts because of what you left out of it; a reference strain drifts because it is alive. Campylobacter has to be replated every three or four days — “it’s like looking after a pet” — and the homopolymer tracts sitting in its capsule and lipooligosaccharide genes flip while you do it, which is how the organism varies its surface and is also, from the bench, your control changing underneath you. One widely used strain carries a plasmid with tetracycline resistance on it that simply drops out under passage, taking any phenotype that depended on it. Seven passages and you may be able to see the difference by eye.

What comes out of this is advice that is unglamorous and entirely correct. Order the strain once. Make a glycerol stock big enough to last years. Work at low passage. Keep two assays whose answers you already know, so that when the answers change you find out rather than publishing.

“God, that sounds like unit tests for microbes.”

It does not rescue the archives, either. A poke around some Salmonella reference strains turned up one or two SNPs between labs — “that was an absolute debacle” — and Typhi CT18, re-sequenced as a technical control, was missing about twenty kilobases of flagellin genes. Ask everyone to post DNA from their type strain and you would not get a point, you would get a cloud.

The cheapest method in the building

Spoligotyping is a presence-and-absence readout of forty-three CRISPR spacers1, which is to say it is CRISPR typing wearing a different name, because in this genus nothing is allowed to be called what it is elsewhere. “It can’t just be called regular things. It always has to have its own mycobacteria in the name.” The variable-number tandem repeat scheme is MIRU-VNTR for exactly the same reason.

It costs under a pound a sample and forty can go through at once, which is why it has outlived every attempt to retire it. It will tell you the lineage. It will not tell you a transmission cluster, and saying so repeatedly in print has not stopped people using it for that. But lineage eight exists because a pattern came up in a routine run carrying none of the spacers anyone recognised, in front of a supervisor who is “just one of those people who has all the spoligotype patterns in her head”. She said it was not a spoligotype they knew. They sequenced it. It was a new lineage.

The tooling around it is going through the same conversion the organism did. Long-read spoligotyping now exists in a tool that was written to do something else entirely and recycled when the original job turned out not to work — “you should never throw away a failed project, you can always recycle it in something else”. The drug-resistance callers are being ported from short reads to Nanopore, which to anyone running them looks like the same program behaving itself, and to anyone writing them is a different program. Somewhere in that port sits the question of what a minority variant is, and the answer differs by tool: ninety per cent of reads at a position, or seventy, or two reads, or five.

Anyway. The pre-genomic method survived because it was cheap enough to run on everything, and running it on everything is what found the lineage nobody knew about. That is not an argument for spoligotyping. It is an argument about sampling, and it applies with more force to the organism that had nobody watching it at all.

The serovars nobody was watching

Ewan Sinclair* ran nine years of population-based surveillance in the Upper River Region of The Gambia, three to four thousand blood cultures a year, taken from every sick child who reached one of nine health facilities. In the early years, twenty to twenty-five per cent of the invasive Salmonella cases were neither Typhimurium nor Enteritidis. Across the last five years it was around sixty. The previous study in the same area, run from 2000 to 2005, had put the atypical serovars at ten to fifteen per cent. The case fatality was concentrated in exactly the group that had been treated as background.

The genomics Ousman Bojang* ran over those isolates was a standard pipeline and none the worse for it: read quality control, assembly with SPAdes, three isolates dropped for having genomes too large to be one organism, screens for virulence and resistance and phage, a core genome alignment, a tree, and Roary for the pangenome. What came out was a gene. The cytolethal distending toxin was known as a Typhi gene, part of what makes a human-restricted serovar dangerous, and it was sitting in more than sixty per cent of the atypical isolates2. The serovars carrying it are not close relatives of each other — the splits between them go back near the base of Salmonella — so this is not one lineage spreading. It is the same gene arriving repeatedly in things that are barely related, by a mechanism nobody has identified, and identifying it means long reads and a careful look at gene synteny rather than anything a short-read pipeline can settle.

The first objection is the counting, and it was put to them in the room. The yearly total is not a quota; it is however many sick children arrived, and that tracks the weather. A heavy wet season brings malaria and respiratory virus with it, so more children come to the facilities and more blood gets drawn. Ewan conceded the mechanism straight away — “if you do more tests, you will find more cases” — and then pointed out that it settles nothing, because the number of cases is not what moved. The proportions are.

The honesty about the rest is worth recording, because it is the part that usually gets smoothed over. Livestock ownership in the study households had not changed over the period. Malnutrition had not changed. HIV prevalence was low and flat. Malaria was falling, and no mechanism was available by which falling malaria would favour one serovar over another. The surveillance was not built to answer the question it was now being asked — “so our data are not ideal”, as Ewan put it — and that was said out loud rather than hidden in a limitations paragraph.

What the finding threatens is the category, and the category is where the money goes. A good deal of non-typhoidal Salmonella vaccine development is aimed at Typhimurium and Enteritidis antigens, which is to say at the two serovars that were receding while this was being recorded.

“If you have Typhi, you have a problem. If you have Typhimurium and Enteritidis, you have a problem. Everything else we’re not bothered about.”

That was offered as the old view and not as the current one. What replaces it is harder to fund, because it amounts to treating the whole species as the thing to watch.

The list TB was not on

TB was not on the World Health Organisation’s priority pathogen list. Not because anyone forgot: the list carries a star at the bottom explaining that TB sits so far beyond everything on it that including it would have swallowed the list. Funding agencies read the list. Nor, Ronan notes, does it appear on the antimicrobial resistance one — “It’s not on the AMR list either, despite it being half a million a year.” The May 2024 revision of the priority list put rifampicin-resistant TB on it as critical.

The TB episodes went out in October 2021, at the one moment in recent decades when TB was not the leading cause of death from a single infectious agent — COVID-19 had displaced it, and the displacement was described at the time as temporary. It was. TB is back at the top of that table. The vaccine is still bacille Calmette-Guérin, and in Camille’s account BCG’s protection is worse the nearer the equator you live, on a theory about prior exposure to environmental mycobacteria; in some places the measured protective efficacy is under five per cent.

The sequencing had not travelled either. In 2021, two countries used genome sequencing routinely for drug-resistance determination. Hybrid capture works, expensively, and the stated worry about it is not technical but distributional — the concern is not creating a two-tier system in which the countries carrying the disease are the ones without the assay. Which is the same shape as the Gambian result: the tools go where the categories say the problem is, and the categories were drawn with the last generation of tools.

The genus problem is worse than any of this and gets less attention. Comparative genomics needs neighbours, and TB’s neighbours are annotated on the basis of TB, so the comparison loops back on itself. An attempt to gather public TB transcriptome data found nothing worth gathering. Whether the organism makes its own B12 or imports it is unresolved, the pathway is present in the genome and apparently unused, and the question of who in a phagosome would be supplying it has no obvious answer.

There is a version of all this that reads as taxonomy being unhelpful, and it was put more directly than that. “It’s a microbe it does what it wants” — you cannot categorise it, and the trouble starts when a label gets treated as a promise about behaviour.

Whoever walked out of that conference had worked on Campylobacter for thirty years and had never seen a capsule. That was true. It stayed true. It was simply never a fact about Campylobacter.

Notes

  1. Kamerbeek J, Schouls L, Kolk A, et al. (1997) Simultaneous detection and strain differentiation of Mycobacterium tuberculosis for diagnosis and epidemiology. Journal of Clinical Microbiology 35(4):907–14. doi:10.1128/jcm.35.4.907-914.1997
  2. Kanteh A, Sesay AK, Alikhan N-F, et al. (2021) Invasive atypical non-typhoidal Salmonella serovars in The Gambia. Microbial Genomics 7(11):000677. doi:10.1099/mgen.0.000677

Outbreaks

Haiti was filling in the paperwork for cholera-free status. Three years with no confirmed clinical case is what the WHO-backed process asks for, and the country was on track to clear it.

Then, in October 2022, with a fuel blockade and gang violence around the capital cutting off clean water for much of the population, cholera came back.

The isolates that reached a sequencer told a clean story. The 2022 strains sat between zero and twenty-five SNPs from the strains behind the 2010 outbreak1. Same lineage. Not an import, not something new arriving — the thing that had been there all along.

Which leaves the question anybody actually needed answered. Where was it for twelve years? That is the one the sequencing cannot touch.

The control nobody pays for

Two stories fit the data equally well. The lineage could have kept circulating in people, at a level low enough and in places thin enough on diagnostics that no case ever got confirmed. Or it could have persisted in the water, because Vibrio cholerae is durable and water is where it lives. Zero to twenty-five SNPs across twelve years is consistent with both.

Even the 2022 isolates were not routine. They exist because a colleague who had worked with Haiti’s national public health laboratory for years deployed into the response and coordinated a subset of cases from several regions of the country. That is why there is a set at all, and why it is a subset.

The comparison set was the harder half. Michelle Kang* and Bryony Jenkins*, who worked the Haiti data at a national public health agency in the United States, needed samples from the quiet years to tell them apart — clinical isolates from 2017 onwards, or environmental samples from anywhere — and there are none. Collecting while nothing is happening is not what surveillance budgets are for. The genomes available for comparison stopped in 2016.

“We don’t have a control for that experiment, if you will.”

Wastewater made this look easier than it is. It worked for SARS-CoV-2, on a precedent set by polio, and both of those are in a sewer because a person put them there. Vibrio is not. It lives in warm water whether or not anybody is ill, so a positive environmental sample proves nothing on its own until you can separate the resident populations from the ones that cause epidemics, and that separation is a bioinformatics problem nobody has closed.

Sometimes the gap is so complete you manufacture the control yourself. Hana Tesfaye’s* PhD on hepatitis A whole genomes began in 2019 and then ran into lockdown, which stopped hepatitis A outbreaks along with everything else. With no cases to sequence, “we had to make our own” — virus spiked into berries and cereals on a bench, because the alternative was a hepatitis A thesis with no hepatitis A in it. Clinical cases turned up once lockdown lifted and the real work could start.

What a SNP distance is claiming

Vibrio cholerae carries two chromosomes, which sounds as though it ought to complicate the analysis and mostly doesn’t. Reference-based SNP calling does not care how many replicons there are; it cares whether the reference is close enough that reads land where they should.

The complication is on the smaller chromosome, which runs to roughly a megabase. It carries the superintegron — about a hundred and twenty-five kilobases of gene cassettes, dense with transposable elements, exactly the sort of neighbourhood where sequence appears and disappears for reasons that have nothing to do with elapsed time. Left in, it generates SNP counts that look like evolution and are really cassette shuffling. So Lyve-SET masks those regions before it counts anything, using a step that is nominally about phage: “it runs a quote-unquote phage-finding program, but really it’s just finding anything that might look like a mobile element or something from a phage”. Anything mobile-looking gets excluded, the SNPs never appear there, and the clock stays roughly honest.

What comes out, then, is a count of differences in the slow-changing part of the genome, between the isolates somebody happened to sequence. Somebody then has to say what the number means. The working rule is ten SNPs for very closely related — near enough to sit in the same transmission chain. Haiti came out at zero to twenty-five across twelve years, which is neither identical nor unrelated, and the reading given was the careful one: related to what had been circulating before, and nothing more than that.

Where ten comes from is a fair question, and the honest answer is that thresholds are conventions. Somebody derived one, for some organism, in some context, and it travelled. The same complaint surfaces in a completely different conversation about sequencing depth — “people randomly make up numbers like, oh, you need 10x or 20x” — where twenty was the working minimum under COVID, for that protocol, that amplicon layout, that base caller. The number outlived all three.

Occasionally somebody does check. Siobhan Reid*, the bacterial bioinformatician at a US state public health laboratory, took long Nanopore reads, which are long enough to span an entire resistance gene, converted the FASTQ to FASTA, ran the raw reads straight through AMRFinderPlus, and compared the calls against assemblies from matched Illumina data. Individual reads still carry errors, but the errors sit in different places on different reads, so with enough of them the noise cancels and the genotype is probably right. The finding held because there was something to hold it against. What coverage you actually need was, inevitably, deferred.

Siobhan also spends more time than anyone intends on separating Klebsiella oxytoca from the species it resembles: Mash against a sketch built from RefSeq representatives, rebuilt at every RefSeq update, with FastANI and skani alongside. The rebuild is the interesting part. The comparison set is not fixed, so the same isolate can be typed one way in March and another in September depending on what got deposited in between. The bioinformatics is provisional. The paperwork is not. “Once you put an organism in an epidemiological report, the epis don’t seem to be able to change that” — and every subsequent analysis in that outbreak has to keep calling the organism what the first report called it.

Resistance runs into the same wall from the other side. Three isolates in the Haiti set carried a distinct profile, genes present or absent rather than point mutations, with gyrA and parC interrogated through ResFinder. Vibrio cholerae is clonal enough that resistance usually arrives with a different strain rather than evolving in place — which is a claim about a population, and only means anything if somebody knows what the population looked like before.

A digression into the samples nobody can name

Clinical metagenomics starts where the ordinary workflow stops. Culture has been tried, the biochemistry and the assays have been tried, nothing came back that helps, so the sample gets sequenced to a depth that would be indefensible if there were any other option. Blood culture, usually. Cerebrospinal fluid where there is one. Skin swabs where there are sores, which run much higher diversity and come with the small compensation that most of what you are sampling is alive rather than fragmented.

These patients tend to be immunocompromised or frail, so polymicrobial infections are common and viral co-infections come along too, all inside a sample that is overwhelmingly human. You can strip the human reads first, and everyone does, but the strip can take the answer with it. The safe habit is to run the analysis twice, on the full data and the dehosted data, and check that nothing important vanished for the crime of sharing a few k-mers with us.

Then the reads go to a classifier and the classifier hands back a taxonomy, and the taxonomy is exactly as good as the thing it compared against. Callum Devlin*, deadpan:

“You can fully trust all taxonomic labels and databases. There’s no errors or mistakes in there.”

Consider what a contaminated reference does. It is the same failure that opens this book — a signal that arrived from somewhere other than the sample — except that it happened once, in somebody else’s laboratory, years ago, and now everybody inherits it. Public assemblies from human-derived isolates routinely contain human sequence, not sitting in an unassigned contig but built into the record — there are Plasmodium falciparum assemblies with a chunk of human DNA in the contigs. Classify k-mers against that and human reads hit falciparum. Early clinical metagenomics duly found malaria and Toxoplasma in an implausible number of people, and the field spent a while wondering about latent infection before noticing the reference was the problem.

The other half of it is the human reference itself. One reference does not capture human diversity, and these samples come from everywhere, so the DNA you most want to remove is the DNA least likely to be represented in the thing you are removing it with. Patients whose ancestry is thin in the reference get worse dehosting and worse taxonomy — a bioinformatics failure with a very unbioinformatic distribution. Pangenome references exist and help. Most prebuilt classification databases still do not use one.

Then there is the flatter failure, which is that the organism is not in there at all. Selection guarantees a share of it: the samples that reach this workflow are the ones ordinary microbiology could not name, which tilts them towards the fungi, the parasites and the environmental oddities nobody has had a reason to sequence. What you have to work with then is reads sharing a little sequence with something in the same neighbourhood, and a phylogeny you hope will place the thing a rank or two above where you wanted it. Whether that space ever fills in is a funding question rather than a method one, and the forecast offered was that money might help and probably would not.

Against all of which, “a healthy person will have 30 pathogens apparently”. Some of that is database noise and some is real and harmless, and separating the two is the same problem again: you would need to know what a person who is fine looks like, sequenced the same way, and nobody is paid to sequence people who are fine. Sequencing every patient in a hospital would find things, most benign, all of which then have to be chased. It is the whole-body-scan problem with organisms in it.

What keeps this honest is that none of it is load-bearing, and Tristin Lindemann* says so plainly. “The doctors and the clinicians aren’t waiting for these results. They’re already treating the patient.” The famous early case — leptospirosis identified by sequencing after everything else had failed2 — cuts both ways. The clinicians were treating empirically all the way through. But the targeted penicillin that leptospira turned out to be susceptible to was started on the strength of the sequencing result, which is what made that diagnosis an actionable one.

There is often no report at the end, either. A report is a legal document and has to come out of an accredited process; this comes out of several rounds of back-and-forth, figures passed across, somebody saying there is a case report that looked like this, can we go further into it. Most of it runs under research use, informing a treatment plan without being part of one.

Lassa, where even the assay is uncertain

Lassa is the version of this where even the assay is uncertain. A biosafety-level-4 haemorrhagic fever virus, mostly rodent-to-human with some evidence of human-to-human spread in clinical settings, and considerably more widespread than the case counts imply — the serology says so. Its diversity is not four tidy clades but a cloud, which is one reason vaccines are hard and the reason usable primers have been so elusive. You cannot design primers around sequence nobody has seen, and nobody has seen it partly because the PCR assays used to find cases may be missing the variants that would show it. Hybrid capture gets around the problem and works, with specificity tunable anywhere from about seventy per cent up to ninety-nine, but it wants equipment and trained people in-country, which is the constraint the whole thing is trying to escape.

So there is a West African amplicon scheme, built with a new primer-design algorithm, sitting untested — untested because the only way to test it is to be standing in Liberia with the primers in hand. Getting reagents into West Africa is expensive and slow enough, through something like ten levels of procurement, that flying a person over with the box can be the quicker route. “Often visiting researchers are like reagent mules”. The primers travel lyophilised, in checked luggage. The amplicons they make are a thousand bases long, which is fine for a clinical sample and probably hopeless for wastewater, where everything arrives already in pieces.

The wastewater work is the part worth watching, because wastewater surveillance is precisely an attempt to build the baseline Haiti did not have: a signal from the quiet years, measured whether or not anyone walks into a clinic. It is funded by the US Department of Agriculture, run out of the National Institutes of Health, alongside a project in Liberia funded by the US Agency for International Development. Two weeks after that conversation went out, USAID stopped existing as an independent agency and what survived of it moved to the State Department.

Cholera has not gone back in its box. It has returned in countries that considered it eliminated, warm water suits Vibrio, and people and goods move in volumes that did not exist in 2010. The tool that could tell a country whether it is genuinely close to elimination is the same tool that could not say where the Haitian lineage spent twelve years, and it will keep not saying, unless somebody pays for sampling in the years when the answer is boring.

The training gap moves slower still. In Callum’s account, genomics was not a formal competency in Canadian medical microbiology residency — exam papers from around 2020 were asking candidates to interpret a pulsed-field gel and decide whether two isolates were related, which is the right question in the wrong technology. Even added tomorrow, residency runs five years, so the first cohort who had to learn it would arrive towards the end of this decade. Public health is further along: getting certified for PulseNet means interpreting this material correctly is part of the job. The clinic is behind, mostly because genomic reporting in hospitals is still ad hoc.

And the labs are thin. The bacterial side of an entire state is one person. Bryony is CLIA-certified as well as being the lab’s bioinformatician, so the same hands run the wet-lab protocol and the analysis. The Lassa work is one person doing wet lab and dry lab and expecting to carry the primers through customs herself. Asked what keeps the work coming, Siobhan’s answer was that “as long as people are getting sick, there’s things to look into”.

Haiti did not get to submit the application.

Cholera-free status is awarded for three years of not finding a case. That is not the same as three years of looking.

Notes

  1. Walters C, Chen J, Stroika S, et al. (2023) Genome sequences from a reemergence of Vibrio cholerae in Haiti, 2022 reveal relatedness to previously circulating strains. Journal of Clinical Microbiology 61(3):e0014223. doi:10.1128/jcm.00142-23
  2. Wilson MR, Naccache SN, Samayoa E, et al. (2014) Actionable diagnosis of neuroleptospirosis by next-generation sequencing. New England Journal of Medicine 370(25):2408-17. doi:10.1056/NEJMoa1401268

When It Was Not a Drill

A hotel lobby in Leamington Spa, 2015, the evening after a day of a microbial bioinformatics hackathon, and the three of us and half a dozen people who would later turn up as guests on this podcast are sitting round a low table playing Pandemic.

The board game. Four diseases in four colours, spreading across a world map a card at a time, and the property that makes it worth a whole evening: it is cooperative. There is no winning it against the person next to you. Either the table coordinates — somebody flying to a city to treat an outbreak somebody else spotted two turns ago, somebody else sitting on cards for a cure nobody has enough of yet — or the board wins, and it wins by doing the same thing over and over in a place nobody was looking.

Nobody at that table thought of it as practice.

Andrew was on jury service for a typical Cambridge crime, assault with a bicycle, when the email came in, asking whether he wanted to join a consortium to sequence coronavirus. It came from Jack Kavanagh*, an office-mate one day a week, who ran pathogen genomics for a national public health agency Wales the other four.

The answer was yes. What the invitation did not mention was that the grant had already been written and submitted, and would be awarded within days. Two weeks after that there were genomes. Samples came through a hospital biorepository that already held the ethics to release leftover diagnostic material, which, as Colm Casey* tells it, is what let the work start almost at once; the consortium’s own approvals arrived much later.

“We started doing everything before even paperwork was done or signed.”

The same fortnight, at another site, Rich Haddon* sent a message on a Thursday evening — “drop everything, I need you to do something, it’s quite important” — and by Monday Alex Pendry* and Tomasz Wilczyński* had bare infrastructure running. The first thing they built was not a pipeline and not a database. It was a user management system, because several hundred people were about to need accounts before they could need anything else.

It was never from nothing

That is the version everybody tells, and it is true, and Rich kept declining to let it stand on its own.

“It wasn’t about standing up an infrastructure from nothing.”

Underneath the fortnight sat five years of CLIMB — the Cloud Infrastructure for Microbial Bioinformatics1, a national cloud built before there was anything national to respond to. Underneath the lab protocol sat the ARTIC network, which existed because of Ebola and Zika. Underneath the trees sat fifteen years of arguing about phylogenetic algorithms in Edinburgh. And underneath all of it, as Jake Radcliffe* set out to a workshop that spring, sat two or three decades of typing schemes — serotyping, phage typing, gels, seven genes and a profile — collapsing into a single method as the cost of sequencing a megabase fell from thousands of dollars to less than one, and staying collapsed for a reason that had nothing to do with resolution: a sequence is a string of letters, and a gel from a pulsed-field tank is a photograph of an afternoon in somebody else’s laboratory.

Take the lab protocol, since it is the piece a fortnight could least have invented. The tiling scheme, as Oliver Wetherby* tells it, came out of a field problem in Brazil: Zika samples arriving at a modal cycle threshold around 36, which is tens of copies, against an Ebola-era method that ran one reaction per fragment and wanted more RNA than anybody had. What replaced it was a commercial assay taken apart from the outside — Thermo Fisher’s AmpliSeq, reverse-engineered on the view that nothing very complicated was going on inside it and the whole thing reduced to a primer-design problem2. The SARS-CoV-2 version came off that machinery on 23 January 2020, days after the third revision of the reference genome was posted. The scheme was new. The apparatus that produced it in a week was four years old.

So the pieces were there, and the assembly was the fast part. What the emergency bought was not build time. It bought permission — to deploy a database that had started life as a sample tracker for faecal metagenomics, to sequence before the paperwork, to skip the year you would normally spend getting it right. Which was the one piece of advice offered to anyone thinking of doing the same. “There’s no point spending a year getting the perfect infrastructure together.”

Days instead of months

The region doing this holds about 900,000 people. Two months in it had produced 1,500 genomes, one for every 600 residents, about as dense as sequencing got anywhere. Twenty-six people were involved at one count, each with a backup, working in pods kept deliberately apart so that one pod going down would not stop the work. The pods were never needed. One person on the project caught the virus, and was still noticing the breathlessness at the top of the stairs many weeks later.

The protocol runs seven or eight hours from sample to sequencer, and its entire purpose is to manufacture an enormous quantity of one specific piece of DNA, which people then carry around a building. Colm Casey’s warning is that “there is a lot of coronavirus nucleic acids moving around the laboratory”, and from that point on every negative control is really a question about your own corridors.

The arithmetic of throughput inverted inside a year. At the peak, hundreds of samples a week made Illumina obvious and a 24-sample Nanopore run absurd — ten of them to clear one week. When case numbers fell and batches shrank to tens, the same arithmetic reversed, and waiting to fill a flow cell became the slowest step in the process. What fixed it was CoronaHiT, which Stephen Hollis* and Colm Casey described to us: a protocol for putting 94 samples on a single Nanopore flow cell at £13 to £18 each rather than about £40, using tagmentation to attach barcodes instead of a ligation reaction3.

That protocol was not designed for this either. It came out of a conversation on a long flight about multiplexing on PacBio, where “every experiment on PacBio is horrifically expensive” and a failed one costs a couple of thousand pounds. The leftover half of a pool went onto a Nanopore flow cell because it was there, unselected for size, and did about as well.

What comes off an amplicon run is not a genome

Multiplex PCR tiles the 30,000-base genome with 98 primer pairs split across two tubes. Two, because adjacent primers in one reaction would preferentially make the short product between them and cover nothing at all. The amplicons are about 400 bases and overlap by up to a hundred, which is what lets you trim the primer off each end and still have the genome covered.

That is the diagram. What comes off the machine is uneven in a way the diagram does not show. Primer pairs differ in efficiency, so coverage can vary by three orders of magnitude across one genome, and lifting the worst-covered regions over threshold is what caps you at 48 to 95 samples on a flow cell. Where a pair fails outright you get a dropout, and a dropout is not a coverage problem. Oliver Wetherby, talking us through the protocol2, is blunt about it: “no amount of coverage will rescue those amplicons, they’re just not there.” Nobody ever gets the far ends of the genome either, and everybody stopped expecting to.

Two things follow, and both bite.

The first is masking. Every read carries synthetic primer sequence at its ends, and if you do not mask it you will call variants in an oligo somebody ordered. At that point an average of six, seven or eight differences separated one genome from the next, so “if you have one false SNP in there, it can cause havoc”.

The second is demultiplexing. Barcodes sit at both ends of every fragment, and the ends of reads are exactly where bases go missing. Any crosstalk and one sample’s reads lift another sample’s empty region over the coverage threshold, and you have called a variant somewhere you had no data at all. An early attempt at 158 samples on one flow cell produced consensus sequences for every one of them and no dependable way to say whose they were.

Which is the whole warning. “A big mistake people make is they just blindly take amplicon data” and then assemble it, or call variants straight off it, as though it came from a genome rather than out of 98 separate reactions with their own opinions.

A digression into a bright pink database

The database holding the UK effort together, which Alex Pendry walked us through, was named after a mask from a Zelda game4 — not, as you might hope, an artefact that averts the apocalypse, but the cursed object that causes it, by pulling the moon down onto the world. Nobody involved appears to have checked. It is either entirely the wrong reference for a national surveillance database or exactly the right one. It was a Django application, and the decision that mattered was that everything went through the API, validation included — so a sample collected in the future, or in 1970, was refused while the person responsible was still looking at the screen.

The test instance was, and this is the correct way to run a test instance, “bright pink so everyone knows it’s the test instance, and we lovingly call it magenta”. The verdict on the code was “small batch artisanal kind of group code”. The distribution pipeline was Nextflow, learned specially for the occasion, with the upstream answer to the first complaint about it being that you should “read the documentation and drink some more coffee”.

And there was Metadata Friday: a hard weekly deadline, a test run of the pipeline mid-morning, then an hour of scrambling to fix whatever it threw up, in front of hundreds of people in a Slack workspace. Miss it and you waited a week. Moving the pipeline to daily runs eventually killed it off, and an attempt at twice-weekly established that “no one really took Metadata Tuesday”.

But the tedium is worth this much space because it was the actual constraint. Compute was not: it was cheap and it was everywhere. Storage was not, because a coronavirus genome is 30 kilobases and a hundred thousand of them with their alignments is nothing. The constraint was metadata arriving from many sources in many shapes, needing to be formalised and validated by hand, and a genome with unusable metadata is not evidence of anything. The uploader most labs actually touched was built by a different group entirely, converted spreadsheets into JSON, and played suspenseful music while it worked, because “consortium work can be quite a drudge”.

What the record does not hold

Everything above is recoverable. The primer scheme has a version number and a changelog. The database has a paper. The consortium has a start date, a closing date and a final report, and somebody could reconstruct the whole of it from the public record without ever speaking to a person who was there.

None of that record holds the part most worth keeping.

Start with the speed, because it is the one thing that had no real precedent. The first genome went up publicly in January 2020, posted by the group that produced it rather than held back for the paper that was obviously sitting in it. The amplicon scheme followed within a fortnight, free, published as a protocol rather than described in a methods section, and corrected in the open by people who had run it and found where it failed. Nobody waited for a signature. There was no pause for a collaboration agreement, a data-sharing framework or a negotiation about author order, partly because there was no time and mostly because for a year and a half almost nobody asked for one.

That is not how any of this normally works. The ordinary version has an embargo on it, and a conversation about who goes first that outlasts the analysis. Groups that had spent a decade competing for the same grants put their methods up before their papers, used each other’s primers, and told each other what had broken. It was never universal and it did not hold — the argument about which database the sequences belonged in ran the entire time, and it is elsewhere in this book. But for a while, in a field not otherwise famous for it, the default flipped from hold to publish, and being scooped stopped being the first thing anybody thought about.

Then the hours, which are harder to write about without embarrassing the people who did them. Almost nobody was doing this instead of their job. They were doing it as well as their job, evenings and weekends, for months, and then for another year. People volunteered for rotas covering work nobody had a job title for, ran flow cells through the night because a sequencer does not care what day it is, and answered the phone at the weekend because a result that arrives on Monday is a result about last week. A great many of them were early in their careers, on contracts shorter than the emergency, doing work that would go out under a consortium’s name rather than their own.

None of it was practised. The protocol was re-cut while it was in use. The advice changed, sometimes inside a fortnight, because the thing being described changed. Almost every method in this chapter was being improved by the people running it at the same time as they were running it, which is not how anybody would choose to work and is what the situation allowed. The willingness to be publicly wrong, quickly, and then corrected by somebody in another country, is the most useful habit the field picked up, and it was learned under conditions nobody would ask for twice.

The bill arrived, and it should be said plainly. Some people did not come back — they left the field, or left research, or spent the two years afterwards recovering from what they had agreed to in the first six months. The version of this story in which everybody was heroic and nothing broke is not the true one, and it is not a kindness to the people who paid for it. Goodwill was a real resource, it was spent, and nobody costed it.

What is worth taking from it is narrower than a moral. It is that the slow parts of this work are slow by convention rather than by necessity. The embargo, the agreement, the year spent getting the infrastructure right before anybody may touch it — all of them turned out to be optional, under sufficient pressure, without the sky falling. That is a finding. Very little has been done with it since.

Where it stands

Take the scaling first — the prediction everybody made and got wrong in the same direction. In October 2020, 75,000 genomes had been through the pipeline and a 100,000-tip tree had just been built, which was called remarkable at the time and “clearly going to need to scale another order of magnitude, unfortunately”. It needed two. By the middle of 2022 the shared database held ten million sequences, and it did not stop there.

Nobody solved this by building trees with ten million tips. They stopped trying. A tree that size is illegible anyway — two thousand tips already “looks like a spiky sea urchin thing” — so global builds were capped at a few thousand sequences, sampled evenly across geography and time, with the background filled in by whichever sequences most resemble the ones you care about. That is the piece that outlived the emergency: not a bigger tree, but machinery for asking a local question of a global database.

In January 2021, the reason given for the primer scheme working so well on this virus was that “all of the genomes present share a very recent common ancestor”, with very little diversity accumulated since. Within twelve months the scheme had been re-cut twice: once in the middle of 2021, and again that December, when Omicron knocked out binding sites the previous version depended on and extra primers had to be spiked in. The mechanism had been named in the same talk — mutation under a primer binding site is one of the listed causes of dropout — so this was less a wrong prediction than a correct one nobody expected to be cashed so soon. A primer scheme is not a published artefact. It is a maintained one, and the maintenance never stopped while the thing was in use.

Also from January 2021: a decentralised network, with its deliberately mixed bag of amplicon, bait capture and metagenomic methods, had produced “what’s been called in the press the best genomic surveillance system in the world”. That was fair comment. It is also over. The consortium closed at the end of March 2023 when the funding ended, and national sequencing continues at a fraction of the volume through routine public health channels. Stood up in a fortnight, stood down on a schedule.

The gap flagged in the summer of 2020 never closed. Only about a fifth of the genomes submitted to the shared consensus database were also reaching the archives that hold raw reads, and the objection at the time was the right one.

“You can’t double-check anything.”

The UK consortium did deposit its reads in bulk, which is a large part of why the read archives hold what they do. Most of the world’s SARS-CoV-2 data exists as consensus sequence only: one pipeline’s interpretation, unrepeatable. The wish list back then already included being able to “re-crunch the entire data set with unified pipelines” that understood mixed positions — which can indicate the direction of a transmission event, and which consensus calling erases. For most of the record that re-crunch is now permanently impossible, and the cause is a submission default.

Sampling metadata went the same way. There is a field recording whether a sequence came from surveillance or from an outbreak investigation, and it is mostly not filled in, so the archive cannot separate a representative survey from twelve sequences off one family — which changes the meaning of any phylogeny built from it, and cannot be repaired now, because the person who uploaded was frequently not the person who collected and both have moved on. Dates are the visible half of the same problem: spreadsheets reformat them, countries order them differently, years get hard-coded, and eighteen months in Hannah Coxwell* was still “fixing dates that come in as Omicron sequences from January 2021”5. Those are the errors that announce themselves. The others are in there too.

And the promise made most often at the start has aged the most awkwardly. Genomics in an outbreak was going to tell a care home or a factory whether it had one transmission chain or several separate introductions: different findings, different interventions. It did that while introductions were still genuinely different lineages, and stopped as any one lineage took over, for the same reason a single false SNP mattered so much. Six or eight differences between genomes cannot separate people whose genomes are identical. What worked spectacularly at national scale was weakest at precisely the scale local public health had been promised.

Two constraints have not moved at all. The entry price for joining a distributed network is still employing your own bioinformatician, which decides who joins one. And the files have become large enough that a slow connection is by itself a reason not to take part. Neither is a bioinformatics problem, and neither attracts a grant — which was the last thing said about the five years of infrastructure that made the fortnight possible. “It’s quite difficult to get a grant to fund that kind of infrastructure work.”

The genomes are all still there. Millions of consensus sequences, in databases that will outlast everybody who filled in the metadata by hand on a Friday morning.

The reads behind most of them are not.

A note on what is missing from this chapter, which is most of it. We recorded about thirty episodes during those years — protocols, variant round-ups, the weekly state of things, people describing what had changed since the previous fortnight — and there is a whole book in them that nobody has written. Almost none of it is here. What survives into these pages is the part that still teaches something now the emergency has gone: how the thing was stood up, why an amplicon run is not a genome, and what it cost. The rest is a record of an enormous amount of work by a very large number of people, and this chapter is a hint at it rather than an account of it.

Notes

  1. Connor TR, Loman NJ, Thompson S, et al. (2016) CLIMB (the Cloud Infrastructure for Microbial Bioinformatics): an online resource for the medical microbiology community. Microbial Genomics 2(9):e000086. doi:10.1099/mgen.0.000086
  2. Quick J, Grubaugh ND, Pullan ST, et al. (2017) Multiplex PCR method for MinION and Illumina sequencing of Zika and other virus genomes directly from clinical samples. Nature Protocols 12:1261-76. doi:10.1038/nprot.2017.066
  3. Baker DJ, Aydin A, Le-Viet T, et al. (2021) CoronaHiT: high-throughput sequencing of SARS-CoV-2 genomes. Genome Medicine 13:21. doi:10.1186/s13073-021-00839-5
  4. Nicholls SM, Poplawski R, Bull MJ, et al. (2021) CLIMB-COVID: continuous integration supporting decentralised sequencing for SARS-CoV-2 genomic surveillance. Genome Biology 22:196. doi:10.1186/s13059-021-02395-y
  5. Omicron was designated a variant of concern on 26 November 2021, so a sequence dated January 2021 cannot be one. It is the date that is wrong, not the assignment, which is what makes it so hard to find.

The Pipeline Wars

Every task Nextflow runs gets a work directory of its own, and every one of those directories gets somewhere between five and ten hidden files in it: the command that ran, the exit code, the environment it ran in. That is a rounding error right up until you cut the work finely enough. One task per genome becomes one for the assembly and one for the annotation, then one for gene prediction underneath that, and nothing about the command you type gets any longer.

“You can submit a million jobs just by accident.”

Eight million small files, give or take, and as many directories, on a shared high-performance computing (HPC) cluster — which is to say a filesystem shared with everyone else in the building. Nobody claimed to have done it, but nobody needed to: that many jobs could overwhelm the scheduler and the filesystem both. It wouldn’t be a bug. The machine just wouldn’t fit the work.

“Maybe HPC is not the right setup then.”

Andrew had recently had 3,800 processors going at once in the cloud, with object storage instead of a shared mount and no inodes to run out of. Do that on a cluster somebody else pays for and a person telephones you and asks you to stop. Do it in the cloud and your credit card melts. That exchange is the entire disagreement in miniature, and it is not really about a programming language. It is about who owns the computer.

Everybody wrote their own, twice

The confession underneath all of this is Charles Lemaire’s*, made in 2022, and it is a confession about years. Bactopia1, his workflow for bacterial genomes, goes back to a predecessor, Staphopia, built for Staphylococcus aureus alone, and that began life as a web form: you uploaded your FASTQ files to somebody’s website, a rack of shell scripts ran on the far end, and hours later you got your results. Then it became Perl. Then Python. Then Python with a workflow library under it. Then it was rewritten again in Nextflow, six years after it started, by which point the field it was written for had moved on. Bactopia was the rewrite after that, because the first question everybody asked was whether it would work for their bug.

By 2024 the commit history kept the score: something like a million lines added and a million taken back out again, with ten thousand or so left standing. Distribution of the predecessor, in the meantime, had been a one-gigabyte tarball containing every binary you needed, hosted on a personal Dropbox, and Charles was sure he’d violated numerous licences doing it. The legal position on that, arrived at some years later, is “let’s hope we’re past the statute of limitations”.

Nobody in the room was smug, because everybody had done a version of it. One of us built a Python contraption as a doctoral student that generated the cluster scripts, copied them across, launched them, and checked afterwards that the output files were bigger than zero bytes. It took months to write, it was hardwired to one scheduler, and it needed a human standing over it between stages. Another noticed, while reading somebody else’s assembler source, a file called step1.done sitting in the output directory, worked out what it was for, thought that is a good idea, and then reimplemented that good idea by hand in one custom pipeline after another. At the Sanger Institute, where Andrew went through three Perl workflow managers of its own, there had been seven or eight written in-house, “all of them equally as poor as the next”. What tipped him over was realising that doing it alone inside one institution wasn’t good enough; the point of a tool the whole community had coalesced around was that it wouldn’t disappear the moment somebody’s PhD finished.

The state of the art being escaped from is worth stating plainly, because it flatters nobody. Nathan Brodsky* brought up Lee’s SNP pipeline, Lyve-SET. While working in Colorado, Nathan had got Lee to install it on a cloud virtual machine so they could distribute an image to other laboratories and spare them the installation. Nathan wasn’t sure he knew anyone else who’d managed it. Lee requested the right of reply, then agreed that it was hard to install.

“I’m done with my defence.”

What we wanted was not a language. It was to stop being the sysadmin. Charles’s account was “I don’t think there’s been a day in my bioinformatics career where I haven’t also acted as a sys administrator”. Our fear was that the field was turning into system administration with a genome attached: anyone who wanted to run an assembler would first have to learn to keep a server alive.

What the unit of work is

A workflow manager takes a description of which tools run in what order, works out what can run at the same time, moves the output of one to where the next expects to find its input, and remembers what already finished so a failed run resumes rather than restarts. The honest version is much shorter: it is an engine, and an engine will not tell you where to drive. You still have to write the thing it runs, and if you write it badly it will execute your bad decisions faithfully, in parallel, on hardware you are paying for.

“Don’t blame Nextflow for your sins.”

The decision it forces on you is how finely to cut the work. One genome, reads to annotation, can be a single task. Or you can split assembly from annotation, then split gene prediction out of annotation, and keep going until every command invocation is its own job. Cut finer and you use more of the machine at once — and you hand the scheduler more to track, and you put more directories on the filesystem, and on a cluster you start paying ten or fifteen seconds of queue latency on a job that runs for ten seconds. Cut coarser, batch things up, and the run gets cheaper and the failures get vaguer: a batch that dies tells you a batch died. This is why nf-core discourages batching, and it is a genuine trade with no right answer: speed against knowing what happened. Nobody offered a way out of it, and there it was left as an open challenge for the community.

The Workflow Description Language — WDL, pronounced widdle, which in Ireland is what you take a small child to do before a long car journey — comes at the same problem from the opposite end, by constraining the person writing the pipeline rather than the person running it. It’s a specification rather than a program: there’s no WDL binary to run, only several engines that implement it.

Jimmy Yoon* put the case for the constraints. Tools are assumed to be in containers before the language gets involved. You are not allowed to make assumptions about where files live, because file paths are precisely the thing that does not survive being moved to another computer. Writing it feels narrow. That is the point: provided the machine underneath can run the containers and has the memory and disk the steps ask for, the engine can put your data on a laptop or in a cloud bucket and nothing you wrote will notice. “In those limitations, it creates more portability.”

But the difference that kept coming back is about something smaller and much more irritating than portability. It is about what the results are filed under.

A Nextflow run is organised by task. Every task gets a directory named after a hash, the sample is a label riding along inside the data, and putting one sample’s ten analyses back together afterwards is a problem for you. “I’ve done an analysis I don’t care what sample it is”. A pipeline author can fix that for their own users — Charles had Bactopia put its results in a fixed structure that its downstream tools know how to read — but somebody has to write that part, one pipeline at a time.

Terra, which is the platform sitting on top of the engine rather than another language, is organised by specimen. As Nathan described it, a workspace holds a table, one row per sample, one column per result, and the cells hold links that resolve to files in cloud storage. When the run finishes, a sample’s results are on that sample’s row, and nobody has to go looking through a directory structure for them.

That reads like a small design difference. It isn’t, for a public health laboratory, because what such a lab produces is a report about a specimen, for someone who asked about that specimen. Nathan’s point was that the people doing that work already read results this way, one specimen per row, because that’s how BioNumerics had shown them PulseNet results for years. Hand them a directory tree instead and you have handed them a second job.

The case against

The argument was eventually held, one of us taking the case against and saying out loud that it was not necessarily his opinion, one taking the case for, one declining to take either and enjoying it. The allocation turned out not to matter much. The complaints that came out are real complaints, whoever was assigned to make them, and the defence ended up conceding some of them.

The case for the defence is not theoretical, and it matters that it is not. A Nextflow pipeline built on a virtual machine at Google, moved onto Google’s batch service by changing a setting, then run months later on Amazon through a commercial front end without the code being touched: “I’ve tried it in three different formats on two different clouds, and it just worked.”

Debugging first. Writing a module is fast — “you can make a Lego piece in 10 minutes, maybe even five minutes” — and the sentence was finished by the person defending the tool. “That can take hours.” The error output narrows a failure down to a module and then stops, and sometimes the executable it is complaining about is not in the pipeline’s repository at all. It is inside a container. So you pull the container, shell into it, find the path, read the script, and then discover you cannot change it, because the container is static and a fix means publishing a new one. The verdict was that “the portability is fine, but it’s obfuscated code”. The person on the fence traced both halves to the same property — “because it’s abstracting, it’s obscuring it away from you” — and even the defence agreed the error reporting needed to get clearer.

Then maintenance. Keep a locally tweaked copy of a shared module and every upstream update becomes a merge you do by eye, forever, followed by a round of testing you also do by eye.

“The fact that you’re editing the code is really your fault.”

The concession that followed was immediate and entirely tactical, and no forked module went upstream that afternoon. The standard being diverged from is, on everybody’s account, a very opinionated one, and plenty of people take an nf-core pipeline as a template, walk off in their own direction, and thereby own every consequence of doing so — which is exactly the position everyone was in before there was a standard, only now with better starting material.

The money is the same argument in a different coat. On a shared cluster you are rationed by a person who can telephone you. In the cloud you are rationed by an invoice that arrives afterwards. Andrew once sent millions of lines of logging into a cloud logging service without noticing, and found out when a thousand-pound bill arrived. Spot instances, spare capacity sold at sixty or seventy per cent off on the understanding that it can be taken back with ninety seconds’ warning, stop saving you anything once a large-memory job has been reclaimed and retried often enough. The upside is real and it is enormous — “millions of nodes if you have a credit card big enough” — and the second half of that sentence is doing a great deal of work.

One bioinformatician, five states

All of this is an argument between people who have a choice.

Some of the people who moved earliest were not chasing scale at all. Jimmy’s own group had spent years handing its viral pipelines to collaborators in West Africa, where the barrier was never processors. It was that a laboratory cannot be asked to buy server racks and then keep them alive on uninterruptible power supplies and diesel generators. Put the compute somewhere else and, provided somebody got the data uploaded first, “your power could go out all day and all night but the compute would go on”. Nothing in the argument about inodes and spot instances reaches that. The binding constraint was the mains.

One state public health laboratory in the American West serves a quarter of a million square kilometres and six hundred thousand people, and when Charles came back on the show in 2024 he was its one bioinformatician. The four other states in its regional consortium had none at all between them. What that lab needed from a workflow manager was not portability across three clouds; it was something an epidemiologist could read rather than a pile of TSV files, which is why Charles’s grant went into reporting and visualisation for Bactopia. The pipeline itself was fine. Andrew had put twenty Salmonella genomes through it and had them back in half an hour, with the mild suspicion that something must have gone wrong because it was too quick.

There is also a step bolted on the front that strips human reads out of bacterial isolate data. Isolates shouldn’t contain any, and Charles was quite confident these didn’t. But the office of privacy and security is busy, the lab had explained that to it many times, and the easiest solution was to scrub every sample regardless, so the lab can now say, truthfully, that everything possible was done.

Which is the answer to the whole debate from the only seat that matters. The one person in that lab does have an opinion about which language won the pipeline wars — he’s been writing Nextflow for years — but the four states around him, with nobody at all, need whichever one somebody else is maintaining.

Where it stands

In 2022 the expectation from the cloud side was convergence. As Jimmy described it, Terra ran WDL on Google, with Galaxy pipelines as a side door for anyone prepared to tinker; Azure had been announced as a partnership but you could not yet put a workspace on it; “Nextflow is on the roadmap at some point”, with some alpha experimentation behind the scenes. Nathan named better harmonisation between the languages as something that would definitely help, and expected tools that looked like competitors to end up complementing one another.

Three years on, the field has coalesced, but not by converging. One tool won widely enough that the others became the things you list at the end of the conversation, and the feature the winner still does not have is the one the other camp already had in 2022. The 2025 discussion runs out on exactly that: one sample, ten analyses, twenty directories, joined back up by hand into the one row a public health scientist can use. The admission arrived with a request attached — “we need it if anyone has solutions” — which is not how a settled argument sounds.

What has moved most is what we are frightened of. In 2022 the fear was being turned into a system administrator. By 2025 the debugging advice from the community is to ask an AI what you did wrong, which is a strange thing for documentation to say, and the reports coming out of the pipeline have a chat prompt on the side of them. Your read-cleaning report goes to a molecular biologist who understands the biology and not the pipeline, they ask the box beside the graphs what the numbers mean, and the answer that comes back is not yours. “It’s whatever the machine is telling them.”

And if a model can write the workflow anyway, the case for this particular workflow language stops being obvious. If the AI generates it, why does it have to be that one — and on those terms nobody needs bioinformaticians, or a podcast about bioinformatics either.

Lee thinks make — the thing that compiles C programs, already sitting on every machine any of us owns — is a perfectly good workflow manager, and reports that “everyone always shoots me down for that”.

The reply he got was about Snakemake.

Notes

  1. Petit RA III, Read TD. (2020) Bactopia: a flexible pipeline for complete analysis of bacterial genomes. mSystems 5(4):e00190-20. doi:10.1128/mSystems.00190-20

Containers, and Other Ways to Stop Suffering

The About page of StaPH-B — the State Public Health Bioinformatics workgroup, getting on for four hundred members by 2022 — credits the founding of the whole thing to a supervisor at the Centers for Disease Control and Prevention who told the first four members to go away. She told them, in Jordan Kowalczyk’s* telling of the page, to “stop bothering her and go talk amongst yourselves”. Nobody had consulted her about it.

“She didn’t even know that she was on there.”

In 2016 those four were more or less the entire population of state public health bioinformaticians in the United States. Each of them was the only one in their building, or one of two, and every question any of them had went to the same federal address, because it was the only address there was. Telling them to email each other instead did not solve a single technical problem. It constituted a profession and the birth of a community.

What that profession eventually built together is a library of Docker images, a few hundred of them, one per tool, and the interesting question about it is not how it works. Docker Hub already existed. BioContainers was already hoovering up everything in Bioconda and turning out an image for each package automatically, and the Galaxy people were doing the same in Singularity. So a group of state laboratories with no spare staff and no spare money went and built their own registry of the same tools, by hand, and were still building it by hand years later. That decision is the chapter.

Fifty of everything

American public health is federated in a way that is genuinely hard to explain to anyone who works under a national health service. Fifty states, fifty procurement processes, fifty sets of rules about what may be installed on which machine, and no shared computer anywhere in the middle of it. There is no cluster where somebody has already built the module and you just load it. Not sharing a single central environment sounds like an IT observation and is in fact the entire problem, because it means the unit of work is one lab, and the same work is being done fifty times, or more if you consider counties within states.

Everyone in this story came from the same place, which is graduate school. Ryan Bautista*, one of the founding four and then at a state public health laboratory, describes the routine: you got hold of a tool by crossing your fingers that the documentation listed the dependencies, working through the dependencies, and discovering that “then you find out those dependencies need dependencies”. A week later, if the versions lined up, “I finally got FastQC running”. Conda and pip made that enormously better and did not fix it, because the thing that stayed broken was the sentence at the end: “it works on my machine”. In research that is a shrug and an apology. In a public health laboratory, where the output of the machine is a line in a report about somebody’s salmonella, it is not an available answer, because the consequences have a real impact on real people.

So two people in two states built containers for the tools the other forty-eight were about to install badly, and that is the whole economic argument. Fifty labs’ worth of installation done twice. It also meant nobody had to be taught how to install Perl libraries, a benefit that got the reception it deserved.

“That is an excellent reason to do anything to avoid teaching people how to use CPAN.”

What the image promises, and what the name does not

The plain definition came from Travis Hudak*, who maintains most of the repository, and it was Docker’s own marketing sentence, near enough word for word — a standard unit of software packaging up code and dependencies so an application runs the same from one computing environment to another. He followed it immediately with the honest version, which is that a container is “an extremely lightweight virtual machine”. Granularity varies. Some hold one tool and its dependencies, SPAdes and nothing else. Some hold an entire pipeline, everything Shovill needs, in one image.

The better explanation came a year later from Siobhan Reid*, at another state public health laboratory, as a ladder. Doing bioinformatics in whatever environment you happen to have is citizen science: you use what is lying around and you can still do good work. A conda environment is a bench with pipettes and a fume hood, where you control rather more. And a container is “doing an experiment on the international space shuttle”, where you control every last thing, at the cost of having had to build the space shuttle. The analogy was conceded to be imperfect about four seconds after it was made, on the grounds that you do not normally do this.

“Just throw away the whole space station and just get a fresh one.”

Here is the part that actually matters, and it is a distinction almost nobody makes out loud. An image is a stack of filesystem layers plus a manifest, and the manifest has a content hash. Change one byte anywhere inside and the hash changes, so the thing you get is a different image by definition. That is real immutability and it is what people mean when they say containers are reproducible. The name is a completely different object. staphb/spades:3.15.3 is a label held in a registry, pointing at a manifest, and the pointer is writable by whoever owns the account. Nothing stops it being aimed at a rebuild six months later. Nothing in your run log will record that it happened, because you asked for a label and the registry handed you whatever that label meant on the day you asked.

Which is why the builds stayed mostly manual for years. That was more work, and Travis thought it was worth it. Docker Hub’s automated builds default to rebuilding the image every time a commit lands on the master branch, and for a web service that default is the feature. For a laboratory that has written the string staphb/spades:3.15.3 into a validated standard operating procedure it is the failure mode itself. So contributions arrived as pull requests, got reviewed, and were built and pushed by a maintainer by hand, so that everything stayed reproducible and stayed under the control of people who could be asked about it. The version of the tool goes after the colon in the tag, so the label carries the one fact you most need to read off it.

By 2022 Jordan had gone a step further and was advising people not to trust the label at all — to pull the image by its hash instead, so that “you know exactly the image that you’re getting every single time”. She also floated locking images so that only one person holds the keys to change them, and, further off, multi-stage builds that run a test dataset through the tool so that a build proves itself before it ships. Every one of those is a mechanism for making a name mean exactly one thing forever, which is a strange amount of engineering to need, and which the rest of us mostly do not bother with, because our images say latest and we have never once been asked to defend a result in front of anyone.

None of that yet explains why an image would sit at an old version of a tool for years, and Ryan’s reason is on the other side of the fence. Upstream releasing something is not, by itself, an argument for rebuilding. It is a proposal that you redo the validation — run your own test set through, show that nothing in the answer moved, write it down — so a release that tidies up a report or makes the tool a little more efficient may not be worth adopting at all. What earns an upgrade, in his account, is a change to the output that helps public health. Old on purpose is a defensible position, and most of the time it is the correct one.

The exception is the interesting bit. Pangolin, which assigned SARS-CoV-2 lineages, shipped a new version or a new model roughly every day or two, and in a pandemic, Ryan pointed out, the newest release was the one decisions were being made on. So Travis, who kept most of the builds under manual control for the sake of stability, wired Pangolin straight into GitHub Actions: a release upstream produced a new image and had it in front of the world about fifteen minutes later. Stability is a value right up until the pathogen is moving faster than your release process.

The ones with a fuse in them

A sealed image is only sealed against the things inside it. Ryan’s war stories were all about containers that were perfectly static and broke anyway. R packages in a report builder that, as far as he could tell, check themselves against the remote version when they run, and fail the job if the two have drifted apart. And a National Center for Biotechnology Information (NCBI) utility that does not merely go stale but has a date compiled into it — the same expiry that has been breaking annotation pipelines elsewhere in this book, arriving here for every image that contains it, at once.

“Oh shoot, oh the date came up, we have to rebuild all the images.”

That is the honest boundary of the whole idea. A container fixes the bytes; it does not fix the world those bytes expect to find. Anything that phones home, checks a clock, or asks a server whether it is still allowed to run has taken your reproducibility guarantee outside and lost it.

Databases are the same problem wearing a different coat, and they are worse because they are enormous. A Kraken database, Travis said, is anywhere from eight gigabytes to a hundred, depending which one you take. The first instinct was to put the small one inside the image, which worked until some users started failing to download it, so the databases came back out and now you bring your own. The software is portable and the reference data is not, and the reference data is where the science actually lives.

The other way to stop suffering

There is a second answer to all of this, which is to write the tool so that there is nothing to install. Sepia1, the read classifier Kees ter Horst* wrote with Lee — Kraken2’s compact hash table, plus a perfect hash function on top, plus a genuine taxonomist’s irritation with how everyone else handles taxonomy — is written in Rust, and Kees’s reason was mostly speed. His everyday scripts are Python, and he reaches for Rust when there are millions of reads and tens of datasets to classify against the clock. Lee had also translated a program of his own into Rust, got a ten or twenty fold speedup for his trouble, and had plainly not got over it.

Kees’s other reason was about reading code rather than running it. He can’t read C or C++, he said, and hopping between header files to follow somebody’s algorithm frustrates him.

“Here you just have one file. That’s where your code is.”

What comes out the far end is a binary that runs. Rust distributes its libraries as crates, and a room with containers on the brain asked the only question available to it — “what do you mean by a crate, is that like a container or something”.

It is not. It is a library, fetched from a central repository like any other. The reason the question is a fair one is that container had already been spoken for twice — once by the shipping metaphor the whole field was living inside, and once, much earlier, by whoever decided that a list was a container — which is roughly the level of naming discipline on offer everywhere here. The tool itself is named after a cuttlefish, as a tribute to Kraken, which is an octopus, and the name nods to the language as well: Kees describes the pigment you can make from a Sepia’s ink sac as rusty coloured — “it’s a humble cephalopod compared to the big kraken”. Nobody who has named a piece of software gets to feel superior about that.

But the single binary does not save you either, and it fails in exactly the place the container does. Sepia’s index over the Genome Taxonomy Database (GTDB) reference set is ninety-eight gigabytes, all of which has to be in memory before a single read gets classified, and the verdict on that was flat and mutual: “so it won’t work on my laptop, it won’t work on your laptop”. Loading it takes about a minute. Classifying a sample afterwards takes about ten seconds — which is why the tool has a batch mode at all, so the minute is paid once rather than once per sample. The expensive part of running the software is not the software. Nothing you can do to the packaging touches that, and the state labs had already found the same wall from the other side when they took the Kraken databases back out of their images.

Nor is a small binary quite as portable as it looks. Rufus Georgiades’s* Deacon, another Rust tool, comes out at two or three megabytes, against more than a gigabyte of conda environment for a competitor written in Python. Then one release went out requiring a newer set of processor instructions that plenty of machines don’t have, because the GitHub machine that happened to build it that day had them. Rufus learned that one by trial and error, and by 2026 Deacon was built for any processor made since about 2013. The binary had carried a piece of the build machine out with it.

Four people, then two hundred

The membership went from four people in 2016 to something near two hundred in 2021 to nearly four hundred a year later. Asked for the proudest achievement, Jordan didn’t put the container library first. She put the Slack workspace first, fifty-odd channels of it, which she checked more often than her email. Siobhan picked the training, and the state bioinformaticians making a Docker image for the first time so that the library could exist at all. That is what the thing was for. The containers are what the conversation produced.

Asked directly why it should continue when other people were building images automatically for everything, Jordan’s answer was validation: images built for laboratories that will one day have to demonstrate, on paper, that the tool which produced a result is the tool they tested. The answer given from our side of the microphone was less institutional and closer to the bone. The automatic builds track the latest of everything, the latest of everything is sometimes broken, and then somebody writes to you about it. “I don’t want the email.”

What was wanted was “more provenance and more conservative set of images, rather than the latest and greatest”, which is a fair description of the difference between software that supports a paper and software that supports a decision. There is also a Python toolkit sitting on top of the containers, which is a package manager wrapped around a container system wrapped around a package manager, and it exists because explaining bind mounts to a microbiologist goes like this:

“So it’s a computer kind of in the computer and you have to do a thing to make it see files outside of the container.”

Two predictions were made in 2022. One was ours: containers and workflow languages were now simply how the work is done, and “if you’re not jumping on that bandwagon you’re going to get left behind”, which has aged into an understatement. The other was Jordan’s, that web platforms would take the command line away from the laboratories without the staff or the compute to run it, and she hedged it even as she made it, on the grounds that things have been threatening the command line for decades and it keeps not dying. She had already split the answer by laboratory, which is how these questions usually end up once somebody frames them as a succession.

And in 2021, asked how anyone should cite the container library in a paper, Travis said that “there is not a paper yet”. The answer wandered briefly through citing the podcast before settling on the GitHub repository and the original developers of the tools. A year later the question had become when, and Jordan’s answer was still not yet, with the pandemic offered in mitigation, entirely reasonably.

What did arrive, in 2023, was a paper about getting bioinformatics implemented in public health at all, which cites the registry as infrastructure and finally puts a number on it: 142 containerised images as of that February2. Which is the ordinary fate of a thing that works. Nobody writes it up, and then everybody cites it for something else.

The supervisor was shown the About page eventually, years after her name went on it, and the report back was that she was “at least a little bit amused”.

Notes

  1. den Bakker HC, Katz LS. (2021) Sepia, a taxonomy oriented read classifier in Rust. Journal of Open Source Software 6(68):3839. doi:10.21105/joss.03839
  2. Libuit KG, et al. (2023) Accelerating bioinformatics implementation in public health. Microbial Genomics 9(7):mgen001051. doi:10.1099/mgen.0.001051

Ontologies, and Other Things Nobody Thanks You For

Nia Llewellyn*, standardising metadata for an interagency project, hit a record saying the sample came from animal layer crumb, and had to work out what that meant. The obvious reading is not a stupid one.

“That must be the crumbs at the bottom of a bag of animal feed.”

That is what went into the standardised field, and the record moved on, and it was wrong. Layer crumb is a feed in its own right — “a special type of feed for chickens that are like layers”, as against broilers, which are a different bird in a different industry doing a different job. The people who knew came back and said so. The field had been correct before it was standardised and incorrect afterwards, and the person who broke it does this for a living and is extremely good at it.

Which is the shape of the entire subject. Not filling in forms. Deciding what a word means when whoever typed it is not in the room, and being wrong in a way that survives, gets copied, and arrives at the far end of an analysis looking like a result. The original field was true, specific and unreadable unless you keep poultry, and that is the ordinary condition of this material rather than the exception — “we see metadata that’s full of jargon”.

The record everybody copied

Take the same failure at the scale of a pandemic. Somebody asked, mid-sentence and half-expecting to be corrected, whether Wuhan-1 — the SARS-CoV-2 reference that every pipeline on earth spent the next three years aligning against — actually has a BioSample record attached to it. The answer at the table was that it did not: a GenBank entry and nothing else.

A BioSample is the structured half of a submission: what the thing was, where it came from, when, from which host, by what method. Without one, a sequence is a string with an accession number and whatever somebody typed into the flat file. Miriam Kessler*, who runs a national food-safety sequencing network in the United States, put it at a rough guess: query the viruses in the International Nucleotide Sequence Database Collaboration (INSDC) and about half come back with no BioSample at all.

What that looks like from the inside is an accusation. One of us had fielded one the day before — “I was accused of not uploading data that I had” — and the data was there, raw reads, BioSample record and consensus genomes, which is more than most submissions carry. It had crossed several countries and several institutions on the way in, and a chain like that gives a field more places to sit than anybody can check.

The omission is not the interesting part. What it did next is. The first record of an outbreak is the one everybody opens when they are working out how to submit their own, and copying the shape of an existing entry is quicker than reading a submission guide in February 2020 — “it’s really easy just to copy that”. So the shape propagated. Sequences arrived with the contextual data smudged onto a flat file, and pipelines got written downstream to scrape it back off again.

In Miriam’s world, foodborne pathogen surveillance had closed this off a decade earlier: one structured format, third-party applications plugged into it, easy to submit and easy to retrieve. She had no reason to think the rest of the world was any different. “I thought that’s how the world worked.” Then the COVID genomes started coming in like that, and “then I realised, actually, in the virus world, that’s not a standard”.

The proposal since is a pathogen data object model — genome data for any pathogen submitted and stored the same way, so that the salmonella standard is not different from the virus standard and nothing downstream has to know which it is looking at. It sounds like something that ought to have been settled in about 2009. The count of ways it is currently possible to get it wrong was put politely: “you would be surprised at how many different ways you can submit a pathogen genome to the INSDC and store metadata in all different locations”.

Nobody is going to be first author on that.

Put your hand over the colours

The demonstration we keep coming back to needs a paper with a phylogenetic tree in it and one hand. Cover the side carrying the coloured tips — the country blips, the host, the date, whatever the legend is for.

“Just put your hand over that bit and just look at this spiky little figure of the tree itself and tell me what does that tell you? It tells you nothing.”

A tree is topology. All of the contrast in it comes from outside the sequence: this clade is poultry and that one is human, this one is 2015 and that one is last month, these two are different wards in one hospital. Take the annotation away and there is no linkage to argue about between genotype and phenotype, nothing about niche, nothing about time. And the annotation cannot be recovered later by being clever.

“You can’t impute it after the fact. You just can’t make it up.”

Everyone agrees with this and almost nobody acts on it. The payoff arrives several people downstream of the cost. Whoever is holding the swab and the form gets nothing out of it — there is “usually no payoff for the one doing the sample collection”, and “the poor sucker who has to fill that form in in the first place, he doesn’t get a look in on what’s going on”. Data scientists get the papers. The person who wrote down which ward it was gets an acknowledgement if they are lucky.

The honest answer to why we personally have skipped it is better than any of the justifications, and it came out of one of us when a guest turned the question around:

“Because I didn’t understand my data and I had to write the paper now. I had to write the paper yesterday.”

Against which sits the only motive anybody offered that survives contact with the work. Everybody says standardisation is the price you pay and talks about ontologies “like they’re bad words”. Nia, who does it all day and has the layer crumb correction to show for it, likes them, and the reason given is not a professional one: “part of it’s probably because I’m nosy and I just want to look at people’s data.”

The same asymmetry runs through the conference programmes. We will happily spend an evening arguing about the nuances of each other’s assemblers and callers, and metadata tooling never comes up, because it is the half of the problem where “they’re not going to put the butts in the seats”. The verdict was that we need to give metadata more street cred, which is precisely the sort of thing a person says about a subject that has none.

What the ID is for

Before any of us would sit still for a definition, one of us went looking for the exit.

“Is there an ontology for hatred of ontologies?”

There is not. There are ontologies for bad words, ontologies at the United States Department of Defence, a proprietary one behind every search you run on Google. There is an ontology for very nearly everything except the feeling the word produces in a working microbiologist.

The feeling is also selective. Anybody who wants rid of ontologies has to give up the Kyoto Encyclopedia of Genes and Genomes, the Clusters of Orthologous Groups and Pfam first, since those are “controlled language for describing gene function” with the word database on the front. We push predicted proteins into them without a second thought, read off “the comings and goings of the different categories”, and call it a result.

The textbook definition, and the one Nia opened with, is that an ontology is a controlled vocabulary where the fields are organised into a hierarchy and there are logical relationships between all the terms. Rui Valadares* always goes for the shorter version: objects, and the relationships between objects — and the relationships can themselves become objects, which is where it stops behaving like a spreadsheet and starts being worth the trouble.

The thing to hold on to is that the unit is not the word. The unit is the ID.

Biscuits are the example Nia uses in talks. A biscuit in the United States is roughly a scone; a biscuit in the United Kingdom is a cookie. Same string, two objects, and no amount of text searching will separate them, because a computer only knows what you tell it — which Rui renders more bluntly as computers being dumb. Give each one an ID and the ambiguity is gone — not resolved by picking a winner, which is what standardising usually means and what happened to the layer crumb, but by admitting there were two things all along. The ontology then carries synonyms, so somebody else’s word for your object can be pointed at the same ID, and relationships, so you can ask a question that crosses two vocabularies at once.

This matters most for words we are certain we agree about. Strain and isolate came up while Nia’s group was responsible for the metadata section of an ISO standard for whole-genome sequencing in food microbiology, and the public repositories turned out not to define them clearly enough to settle it. SNP is worse, and Rui’s version of why is that a SNP only means something exact once you know how it was called. It is defined by the set of tools and parameters that produced it: ask a human geneticist and a microbial bioinformatician for a definition and you get two different answers, and ask two microbial labs and you can still get two. Both will write SNP in the column header.

Get that right and the payoff is real. Nia’s example is the antimicrobial resistance ontology underneath the Comprehensive Antibiotic Resistance Database (CARD),1 which ties antibiotics to their mechanisms of action to the genes, and is why the tool on top of it can do more than string-match gene names.

Get it wrong and you have not fixed the problem, you have upgraded it. Structured silos are still silos. Ontologies are built by people, Nia warned, and people have jobs — animal health, human health, agriculture, environmental research — and an axiom that is obviously true in one of those can be quietly false in another, so two ontologies written independently can disagree about what is a subclass of what. Run a reasoner across both and the clashes come out. There is a community, the OBO Foundry, with term reuse among its key practices for this reason, and it still happens. The cleanest illustration came from the two people in the conversation who each build these things: Nia had a genomic epidemiology ontology, Rui a typing ontology, and “those won’t even work together because we designed them differently”. The interoperability people are not interoperable.

Which leaves the free text everybody actually has. Nia’s group built one of the mapping tools, which tokenises short records and converts them to standardised terms, and such tools are worth using, but in Rui’s account a good run is seventy to eighty per cent and some of the rest is done by hand. The only real fix is upstream, and the analogy reached for is old and correct: bring a statistician a finished experiment and, as Ronald Fisher had it, all that is left to do is the post mortem — “the only thing I can do is an autopsy on the data”.

The bats in the drawers

The SARS-CoV-2 specification — which two of us worked on, and which this chapter is therefore not a disinterested account of — had to describe considerably more than a nose swab, which is where it gets enjoyable. Sewage, door handles, air ventilation shafts. Pets, livestock, wild animals. Breast milk, saliva, blood — because the case for a less invasive diagnostic test only exists if somebody has already recorded viral load in the thing you would rather swab.

And then the use case that sticks, for which Miriam had the specifics: a global consortium of museum curators, building out a spec for broad-based coronavirus surveillance in museum specimens. Bats in drawers. Decades of reservoir sampling already done, by people who collected for entirely different reasons and labelled everything anyway.

The specification’s reference guide has something called jitter, too, which is the same craft from the other direction. A collection date can identify a patient in a small enough population, so one of the mitigations the guidance offers is to shift it — add a day, subtract a couple. A value deliberately slightly wrong, so that it is safer to share.

The museum drawers are the argument in miniature. The specimens exist. The sequencing is affordable. The only thing between a cabinet of preserved bats and a reservoir survey is whether anybody agreed on a field that can say what the thing in the drawer is, and whether the curator who wrote the label in 1961 used the same word for it that you do.

Where it stands

When that specification went out in August 20202, there were about 12,000 SARS-CoV-2 genomes in the INSDC and 75,000 in GISAID, sixteen UK sequencing centres feeding one central repository, and the Quadram Institute’s share of that was around 1,600 genomes. Those numbers are the reason the spec exists and they date it better than anything else in it. The prediction attached to them held in East Anglia, England: local cases had fallen from about fifty a day at the peak to two or three a week, and the flat expectation was that it would climb again “maybe in a few weeks or a few months”. It did.

The specification itself travelled. By Alasdair MacLennan’s* account in 2020, it had been adopted by the Canadian sequencing effort and by labs in the American network, was being circulated to African labs and was being considered in Australia — built, as Nia described it, on the existing standards rather than as a replacement for them, on the grounds that it “isn’t meant to be a panacea”. Miriam’s case for it was modest and has lasted: if everyone agrees a minimum, “then we can really have a global interoperable surveillance network”.

By the time we were sitting in Vancouver in 2023, the argument had moved from the spreadsheet to the repository, which is a promotion. The interesting development is institutional rather than technical. There are now two organisations doing this — the Public Health Alliance for Genomic Epidemiology and the Global Microbial Identifier. Miriam, unsure what the future of the second one was, pointed to its intergovernmental presence — the World Health Organisation, the Food and Agriculture Organisation, ministries of health — and suggested it could take the first one’s standards and implement them, “which I think is a much harder part than just creating the standard”. We liked that as a division of labour: one writes the specifications, code and papers, the other has the contacts. Whether one body should simply become a subgroup of the other was floated at the table and got the only sensible reply, which was that it would be a disaster, followed immediately by no comment from the government people present.

That conference took four years to happen — “it takes four years to think about it, and then a few months to organise it”, in the account of Henry Liao*, who put it on with Nia. What it did differently was treat data sharing as a social problem with a technical component rather than the reverse: First Nations and marginalised community participants on the programme, a gender balance bioinformatics does not usually manage, an explicit refusal of one solution for everybody. A major First Nations conference running across town the same week made that harder, since group after group was already going to it. That work, Henry was clear, was Nia’s, and he would take none of the credit.

The claim from late 2020 that has aged most sharply is Rui’s: that the genomic data is now the easy part and the metadata is the hard part, so metadata may be the next frontier. That was said when sequencing a genome had just stopped being a technical marvel, and it was right, and it is right in the least satisfying way available: three years later the frontier was still Miriam explaining that a virus submission and a salmonella submission should look the same. Nia described ontologies in the same conversation as slow in moving into public health.

Somebody in a poultry operation wrote down exactly what the sample was, in the only words that were exactly right, and an expert tidied them into something plausible and wrong, and the only reason anybody knows is that the person who wrote it argued back.

Everything downstream of that field — the tree, the cluster, the outbreak call, the paper — depends on who won.

Nobody has ever been thanked for winning it.

Notes

  1. McArthur AG, Waglechner N, Nizam F, et al. (2013) The comprehensive antibiotic resistance database. Antimicrobial Agents and Chemotherapy 57(7):3348-57. doi:10.1128/AAC.00419-13
  2. Griffiths EJ, Timme RE, Page AJ, Alikhan N-F, Fornika D, et al. (2020) The PHA4GE SARS-CoV-2 contextual data specification for open genomic epidemiology. Preprints.org. doi:10.20944/preprints202008.0220.v1. Published as Griffiths EJ, et al. (2022) GigaScience 11:giac003. doi:10.1093/gigascience/giac003

Who Gets to Sequence

A freezer in the MRC Unit in The Gambia, January 2020, holding aliquots of samples that are of no use to anybody.

They are there because they have to be. Any sample leaving the country has to leave a portion of itself behind, as a condition of ethics approval, and the portion that stays is not linked to any of the metadata that would make it worth keeping. So it sits in a freezer, taking up space and electricity, and the person describing it is entirely clear-eyed about the cost: “we have a lot of samples that are practically useless, but we have to keep an aliquot back, and it’s not linked to any of the metadata, which is a pain.”

Then, immediately, the other half of it.

“But this is a regulation, yeah.”

Mansur Kamara* put the reasoning plainly. Governments have worked out what the arrangement used to be, and are changing it: “that sample doesn’t belong to research scientists it belongs to the country so they can come here and seize our freezers”. The ambition underneath is not paperwork. It is to sequence in country, so that you do not have to “ship it off to Sanger or to Broad or wherever”.

The rule everybody has and nobody notices

The thing that makes that room worth reporting is what happened next. Samina Qureshi*, working in animal and plant health, agreed with all of it and then pointed out that her own institution works the same way. Anything funded by the government — “which is 95% of our work” — belongs to the government rather than to the scientists, and if people leave, it stays. Permission to do anything else goes through policy people.

Two laboratories, one in The Gambia and one in the UK, describing versions of the same principle: the samples aren’t the scientists’ to dispose of. Samina took hers to be standard for government work. What differs is not the principle but the consequence. If your samples belong to your government and your government has three sequencers, the rule is an administrative step. If your samples belong to your government and the nearest machine that can read them is in another hemisphere, the rule is the whole question, because it decides whether the work happens where the samples are or where the instruments are.

That is the shape of nearly everything in this chapter. Not one rule for rich countries and another for poor ones. The same rule, landing differently, because the arrangements around the technology never got cheap at the same rate the technology did.

What capacity actually means

Sequencing costs collapsed by orders of magnitude across the years this book covers, which is the fact everybody quotes. The costs that did not collapse are the ones nobody puts on a slide: a person who can run the analysis, a machine to run it on, a network connection that can move the data, and somebody to keep all three alive next year.

Alasdair MacLennan*, who helped establish the alliance we’ll come to shortly, is careful about where the shortfall sits. In large surveillance programmes such as PulseNet, he said in the summer of 2020, the standardisation and deployment of wet-lab methods was “pretty well understood and well defined”. What genomics added was a bioinformatics and data-management load that many public health laboratories had no way to carry:

“These shortfalls in infrastructure and workforce capacity are both critical and widespread in public health, regardless of where they are in the world.”

Regardless of where they are in the world is doing real work in that sentence. This is not a description of somewhere else. “Most labs just don’t have access to a dedicated team of bioinformaticians”, and they do not have system administrators, and they do not necessarily have cloud resources — and that is as true of a state laboratory in the American midwest as it is of a national laboratory in west Africa. The containers chapter of this book is what that looks like when fifty American laboratories each face it alone.

Then the software itself, which offers a choice between two things and nothing in between. Tools are either “really customised and open source”, meaning somebody has to be able to drive them, or they are proprietary, meaning somebody has to be able to pay.

“There’s really no middle of the road.”

A laboratory with neither a bioinformatician nor a budget is not underserved by this market. It is invisible to it.

An alliance named after a virus

What came out of that was PHA4GE, the Public Health Alliance for Genomic Epidemiology, which is pronounced phage, because a room full of microbiologists was never going to let that go.

It was modelled on something that already existed — “the Global Alliance for Genomics and Health on the human genomic side”, which had grown up around the realisation that human genomics needed common standards. Alasdair and his colleagues first asked it whether it would take on microbial working groups; it was supportive, but encouraged them to set up on their own. The public health side had groups thinking about standards and openness in patches, but “there really wasn’t any traction or forward motion on a lot of them”.

The detail worth stopping on is where they put it. The secretariat went to the South African National Bioinformatics Institute in Cape Town. Not Atlanta, not Cambridge, not Geneva. For an organisation whose entire premise is that capacity is unevenly distributed, the address is the argument, and it is the kind of decision that is easy to describe as symbolic right up until you notice how few organisations make it.

It launched at a Grand Challenges meeting in Addis Ababa in late 2019 and planned to spin up its working groups over the following spring. Then the pandemic arrived. By the time we recorded with Alasdair that summer, three or four of its eight working groups had formed and the rest were still coalescing. Their names are a fair inventory of what the field thought it was missing: data structures, infrastructure, bioinformatic pipelines and data visualisation, training and workforce development, public sequence repositories and ethical data sharing among them.

The one worth watching is infrastructure, because one of its first tangible outputs wasn’t software. It was a survey tool “to reach out to the various public health entities to really get a sense of what their bioinformatic capacity, what their bioinformatic resource needs are” — using SARS-CoV-2 as the basis, since that was what everybody was doing, but aiming at the general question of “global access to cloud services, how much technical capacity there is out there”.

Read that twice. In 2020, in the middle of a pandemic being fought with genome sequencing, the alliance set up to improve public health bioinformatics was still having to ask who could do it. The gap wasn’t only in capability; it was in knowing where the capability was.

The part that cannot be solved with a standard

None of this is only about money, which is the comfortable version of the story.

Some of it is law. Genetic resources are governed by an international access-and-benefit-sharing framework1 that different signatories read very differently, so material collected in one country may never reach anybody else’s laboratory — a constraint that shows up in this book’s taxonomy argument as an obstacle to naming things, and shows up here as holes in the comparison sets.

The argument has since moved beyond physical samples. In 2022 the parties to the Convention on Biological Diversity agreed to set up a multilateral benefit-sharing mechanism for digital sequence information, and in 2024 they agreed how it would run, including a global fund, the Cali Fund. The Pandemic Agreement adopted by the World Health Assembly in May 2025 left its annex on pathogen access and benefit-sharing to be negotiated afterwards. Putting a sequence where anyone can reuse it doesn’t, by itself, settle who should benefit.

Some of it is reference data. Nearly every method described in this book compares a new genome to genomes already deposited, and what is already deposited is whatever somebody had the capacity to deposit. A database is a record of who was sequencing, not of what exists. It shows up even in quality control. A SARS-CoV-2 genome carrying private variants, ones nobody else has deposited, can get flagged as poor sequence when it may simply have come from an under-sequenced part of the world. At a workshop in 2021 we told people to take those flags with a pinch of salt.

Some of it is delivery. In early 2021 Andrew’s group had a collaboration with Zimbabwe, at a point when so many borders were closed that shipping anything was hard. The laptops and a Nanopore device could travel at room temperature and went out without trouble. The reagents travelled on dry ice, which lasts a few days, and they spent two weeks in a warehouse at Stansted Airport. Thousands of pounds’ worth were destroyed, and that was the urgent shipment, the one Andrew’s group had paid extra for.

And some of it is that the tooling is designed by people whose conditions are not universal. A pipeline that assumes a fast link to a public archive, or a GPU in the cloud, is making an assumption about electricity, about bandwidth pricing, about whether the connection stays up for as long as the upload needs. Those assumptions are rarely written down in the software, which makes them hard to argue with.

Where it stands

The alliance still exists, which for a volunteer standards body founded months before a pandemic is not nothing. Its SARS-CoV-2 contextual data specification2 went out in 2020, was adopted in Canada and the United States, considered in Australia and circulated to African laboratories, and is described elsewhere in this book by some of the people who worked on it, two of whom are writing this. It was never a demand to share everything. The idea was to get the data into a standard shape from the start and then decide what went to a public repository, what went to trusted partners and what stayed in the laboratory for local analysis. Agreeing on the columns didn’t mean agreeing to hand over every row, which matters to anybody whose samples belong to their government.

The part it was set up for is harder to mark. The survey of who can do this work was a first attempt at a question that doesn’t stay answered, because capacity is people and funding, and both move. Meanwhile the funding that supported a good deal of the work in low- and middle-income settings has been reorganised or withdrawn, on timescales that have nothing to do with whether the work was finished.

Our own recent experiments get round a small part of it. The browser tools we were building in 2026 compile existing programs to WebAssembly, so the page downloads once and then runs on the user’s own machine; the reads never leave it, and nobody has to push a large FASTQ file up a poor connection. It doesn’t give a laptop more memory than it has, or supply somebody to interpret the result.

The technology kept getting cheaper throughout. That was never the constraint, and treating it as though it were is the most comfortable mistake available, because it is the one that resolves itself while you wait.

A machine that reads DNA now costs less than a decent car. A person who can tell you what it said still costs a career.

Notes

  1. The Nagoya Protocol on Access to Genetic Resources and the Fair and Equitable Sharing of Benefits Arising from their Utilization to the Convention on Biological Diversity, 2010.
  2. Griffiths EJ, Timme RE, Page AJ, Alikhan N-F, Fornika D, et al. (2020) The PHA4GE SARS-CoV-2 contextual data specification for open genomic epidemiology. Preprints.org. doi:10.20944/preprints202008.0220.v1. Published as Griffiths EJ, et al. (2022) GigaScience 11:giac003. doi:10.1093/gigascience/giac003

Publishing, Reviewing, Shouting Into the Void

“What’s something that is unownable, you cannot own this thing, and I’ll make an NFT out of that thing so I can own it.”

The answer came quickly, because a court had already supplied it. Myriad Genetics had held patents on BRCA1 and BRCA2, and in 2013 the Supreme Court of the United States struck them down1 and wrote out the reasoning at length: a gene you found sitting there in a person is not something you invented, and it is not yours.

So, on a day off in lockdown, Lee made a picture of the exons of BRCA1 in Apollo, took ribbon and cartoon renderings of the protein from the Protein Data Bank, found a double helix for the background, pasted in a few lines of the court’s decision, and minted the resulting picture as a non-fungible token. Thirty dollars, plus eight more in virtual gas, because moving anything across Ethereum costs money and somebody’s computer has to be paid for its trouble.

Then nothing happened. Nobody bought it, nobody cared, and the relief was genuine. Lee, a US federal employee, remembered the limit on gifts as about ten dollars, and anything from a foreign government as worse, and the other two people on that call were paid by the British government.

That is the joke, and it is a good joke, and it is not the part that lasted. What lasted is that the exercise worked exactly as advertised. There is now a permanent, public, timestamped record, which nobody can quietly edit, of one person having done one specific thing on one specific afternoon. Which is more than our own field manages for most of the work that keeps it running.

The ledger nobody keeps

It took ten minutes of explaining the blockchain before we had reinvented the accession number. If a public ledger can record a transaction involving a picture, it can record a genome deposit, so mint a token per assembly, staple it to the paper, and retire the central archives — the National Center for Biotechnology Information (NCBI), the European Nucleotide Archive and the DNA Data Bank of Japan, all liberated from the business of hosting everybody’s data forever. The objection came fast, and from inside the same conversation: one of us was, on reflection, “horrified at the prospect of destroying NCBI”.

There was a more basic problem, too. The token could hold an accession and a checksum, but the genome file would still need to live somewhere else. We hadn’t replaced the archive; we’d added a transaction fee to the accession number. A database was usually better than a blockchain, and tens of dollars per transaction times every genome anybody deposited was a sum nobody wanted written down.

Somebody asked what would happen about retractions.

“It would never happen, because science does everything perfectly and there’s never any mistakes.”

From there it went where these things go, which is to a currency. BioBucks, or GenomeCoin, awarded for the things that no existing system counts: answering somebody’s bioinformatics question, helping a student, reading a paper, reviewing one. Then the immediate collapse into what such a number would actually be used for, which is to be printed on a form for a person who has no idea what it means.

“I don’t know what that means, but it’s a big number, so it has to be good.”

Somewhere among the island-buying and the double helix in place of the Bitcoin B, one line went past without anybody stopping on it: “I don’t think getting paid for reviews is a good thing, but getting some recognition would be good.”

That is the whole problem stated in one sentence, as an aside, in a joke episode about cryptocurrency. Not all of us would have signed the first half of it. One of us answered review requests from commercial publishers with an offer to do the work at the full economic cost of his time, and nobody had ever taken him up on it.

Because the accounting we do have is for papers, and papers are the smallest part of it. The version of this that hurts is the ordinary one — “someone will spend years working on a paper and then silently drop it into the ether”. It surfaces in a search index somewhere, possibly in a decent journal. Nobody pulls out the finding, nobody says why it matters, and the result the community actually needed is sitting on page 39 of 50. Work inside a government agency has the same shape with the credit removed at source — it goes up under the agency’s name, and “you’re a faceless bureaucrat”.

Academia manages its own version. For the many years one of us worked at the Wellcome Sanger Institute, its staff directory listed only the principal investigators. That changed only in 2018, and everybody got a photograph.

None of the fixes proposed for this are interesting, which is the point of them. Get onto your own organisation’s staff page, because a search engine trusts an institution with a physical address more than it trusts you. Claim your Google Scholar profile, because an unclaimed one drifts down the page and a claimed one takes five minutes. Merge the preprint with the published version so the citations add up to one number instead of two. Register an Open Researcher and Contributor ID (ORCID), which is pronounced orchid, spelled without the H, making it an orchid ID, a double ID, and we did not resolve this. Put your name inside your own scripts so that anyone running your tool with a help flag finds out who wrote it. Get a GitHub account that isn’t called something like “lighthouse seagull 52” as this is your portfolio of professional work.

And then there is Publons, which existed to record the reviews. You told it what you had reviewed, or the journal told it for you, and it could show that you had done eight for one journal and a couple for another — exactly the fodder grant forms demand. The objection arrived in the same breath as the recommendation. You are “giving your data away to a huge commercial company worth billions so that they can then use it and sell it on back to your organisation in a different format”. Both things were true, and most of us had signed up anyway, because “we do it for free, it’s a community service” and it ought to be counted somewhere. It didn’t have to be Publons, either: where the journal supported it, the review could go straight onto your ORCID record instead.

What you read, and in what order

Somebody admitted, on air, to taking two days over a review and being embarrassed about it. Somebody else said “five minutes if it’s terrible”. The gap between those numbers got a title proposed on the spot, as a riff on a romantic comedy — “how to review a paper in 10 minutes, great title”. It got made, and by the time it did, ten minutes had become ten days.

We didn’t all read in the same order. Some of us went straight through on the first pass. The alternative is worth having, because it isn’t about reading faster. It’s about reading in an order that lets you stop.

For the first pass, start with the last paragraph of the introduction, the one that says “here we aim to”. If that makes sense, leave the rest of the introduction until later. Then go to the figures, before any of the body text, and get the gist from the pictures. Then skim the results. Then, for anything the results actually depend on, go into the methods. Half an hour, maybe an hour, and then put it down and go away and have lunch, because the first pass is not the review. The first pass exists to find out what the paper is trying to do, which the paper itself frequently declines to tell you in the order you need it.

The second pass is the review, and it runs in the order that matters rather than the order the manuscript is printed in.

“It’s a house of cards, any publication.”

Methods and results are the base layer. If those are wrong then the discussion is wrong, the inferences drawn from them are wrong, and there is no point tidying anybody’s grammar or telling them which citation is missing from an introduction that is describing an experiment that did not work. You say the base layer is unsound, you say why, and you stop. That isn’t rudeness. It is the only version of the job that is honest about your own time.

For our kind of paper, the base layer is checkable in ways a wet-lab reviewer would find strange. You cannot repeat somebody’s sequencing run, so you go after the proxies. Are there accession numbers, and do they resolve to anything. Is there code, and where the paper is a piece of software the position hardened into a rule — “if they don’t provide the code, that is an immediate reject from me”. Does the code do what the paper says it does, or is it “a 300 line bash script and it’s full of just print statements” being offered as a pipeline, or another tool run with custom parameters and the output reshaped, offered as a new method. Are the tools even pointed at the right data: a long-read assembler run on short reads, an assembler nobody has touched in twenty years, a platform’s error profile ignored entirely. A methods section is the one part of a manuscript with nowhere to hide.

Then the parts that keep a review honest. If there is a bench section you cannot assess, say so, defer to whoever can, and stop agonising over it. Read the journal’s actual scope, because at PLOS ONE technical soundness is the whole test and novelty explicitly is not, while another journal will make you write a paragraph on novelty you’d otherwise never have thought about. And do not review a paper into the paper you would have written — “it’s my way or the highway. That’s not your role as a reviewer” — which is the most common failure of all, and one you get to watch happening, because the other reviewers’ comments arrive with the decision and there is always one.

Which leaves writing back. On a few occasions the note to the editor has said that “reviewer 3 is a complete idiot and does not understand the field”, has asked for work nobody needs, and should be ignored. That is arguing the author’s corner against a colleague neither of you will ever meet, and it is occasionally the most useful thing a reviewer does.

Occasionally it goes the other way, and this is the good day. You reach page 21, find the actual result of the paper sitting there unmentioned in the abstract, and write back to tell the author that the work is better than they realised. Meanwhile, somewhere else entirely, “someone went and rebranded a heat map”, called it the quilt plot, and got it published.

Not a murder mystery

The reason a review takes ten days rather than ten minutes is that most papers are written to be read twice. Everything is arranged as a reveal: here is a data set, here are the figures of the analysis, and now, in the discussion, the thing we found. Half those figures are not driving the point at all. “Scientific writers love to have foreshadowing in the work”, and the reader has to hold all of it in their head to work out afterwards which parts mattered — “it’s like a film you have to watch twice because you don’t catch everything the first time around”.

The best correction came second-hand, from a food microbiologist with a Schwarzenegger accent, telling his own students:

“This is not a murder mystery. Don’t make me wonder what it is. Just tell me.”

It is difficult advice to take, because school spent years teaching the opposite. A paper is not a story. It is a technical document, its job is to hand the information over, and you may “spoil it as much as you like”.

Which lands back where the first conversation started, from the other side. The person who buries a finding on page 39 and the person who drops a paper silently into the ether are the same person in two places. And the reviewer who digs it out is doing the promotional work the author skipped, for free, anonymously, in the evening.

After Twitter changed hands

In late 2022, Twitter changed hands and a chunk of scientific conversation went looking for somewhere else to be. Nabil-Fareed and a collaborator bought a domain, spun up a virtual server and had an instance of Mastodon running in a couple of days — a Ruby and JavaScript web application with a Postgres database, a Redis cache and a background job runner, which if you have ever built a website is not exotic. The one genuinely hard part was none of those. It was sending email — “the most expensive thing is actually getting a mail server that works” — because the whole internet is now defended against spam, and a password reset from a small honest server looks exactly like a password reset from a dishonest one. The expectation was fifty people, maybe a hundred, a place for microbial genomics types to post memes.

“Let’s just try it out. What’s the worst that could happen?”

Two thousand accounts inside a fortnight, arriving at two or three hundred a day. A Nobel laureate. Journals. Historians and pharmacologists and immunologists, none of them in the old feed. The web server processed about one and a half million background jobs in a week, because a federated network spends most of its effort replicating other servers’ content whether or not your own users are doing anything. It cost around a hundred pounds per thousand users per month, which is a “substantial amount for one person to carry”, and the software has no advertising mechanism even if you wanted one.

The technical prediction was correct and boring. A decentralised network is slow, because every request to another server is a request that might not come back; you cannot index content you do not hold, so search will always be bad; and therefore Mastodon “won’t replace Twitter as a one-for-one replacement but it might fit certain use cases”. That has held. The first complaint was that you couldn’t find anybody, and it still is, a partial full-text search having arrived the following year and not settled it.

The social prediction was put as a straight question:

“Is this the end of Twitter, or is it just a temporary flash in a pan and we’ll all go back?”

Mostly the second, and not in the way anybody expected. Sign-ups peaked within weeks and most of those accounts went quiet over the following year. Twitter became X in 2023 and kept a great many of the people who had loudly left it, for a time. The ones who did leave for good largely did not stay on a federated network either — they ended up on Bluesky, a platform run by a single company, which is the precise architecture the migration was supposed to be an argument against. Hannah Coxwell* called her own move “a little bit of bet hedging” and had no plans to leave the old place, although that later changed.

The part that dated best is the invoice. A great deal of scientific conversation was sitting on one person’s virtual server, paid for monthly by that person, and moderated by that person too, blocking dodgy servers by hand, because federation makes every badly-run corner of the internet a neighbour. It is the whole story of who pays for scholarly infrastructure, at speed and in miniature.

And Publons is gone. Clarivate retired it in 2022 and moved the profiles into Web of Science, so the record most of us had been keeping of our unpaid reviewing now lives inside the commercial product that sells the metrics back to the institutions employing the reviewers. The objection made on air in 2021 did not age. It simply completed. The ORCID route mentioned on the same call is still there, for the journals that support it.

The token for the picture of BRCA1 is still on the blockchain, the record of it cannot be altered, and it cost thirty-eight dollars. The reviews were free.

Notes

  1. Association for Molecular Pathology v. Myriad Genetics, Inc., 569 U.S. 576 (2013).

Machines That Write Code

“Mine has gotten rid of all authorship and it’s just like, Lee, sorry you’re gone.”

June 2023, live on air, no rehearsal. Lee pulled a small Perl script out of his own repo — the kind that takes a non-standard FASTA file with the entries separated by pipes and turns it into something a normal tool will read — and we fed it to a large language model to see whether it would come back as Python. Then we fed the Python back and asked for Perl again, to close the circle.

It came back with argparse and Biopython. One copy kept the author’s comments at the top. The other had quietly deleted them. Both looked right, and nobody ran either of them on air. The round trip finished with a disclaimer: the conversion had been done to the best of its ability, and there might still be slight differences requiring adaptation. It had learned to cover itself before it had learned to be right.

Nobody was alarmed, exactly. “It might put us out of business if it’s writing code this good.”

That was the last moment any of this was small enough to check by eye, and none of us noticed. Everything we settled over the next three years — how to hire, how to examine students, what a tool was even for, whether a model can be an author — rested on one assumption nobody argued with.

The machine is fine as long as a human reads what comes out. Nobody asked how much a human will read.

The part everybody agreed on

The agreement was about boilerplate. A column of dates written six different ways, converted to ISO format. The pandas syntax you look up every single time because you have never used it often enough to keep it. Unit tests, which are a typing exercise rather than a thinking one. Ten dollars a month to GitHub for Copilot, and twenty or thirty per cent off your coding time straight away.

The caveat was just as unanimous, and it was not about correctness in general but about the specific way these things fail. The model “has a wonderful habit of lying when it doesn’t know”, and it lies in exactly the register it uses for everything else. It cannot flag its own uncertainty because it has never been allowed to. “The machine is never allowed to say, oh, I don’t feel confident about this.”

It is worst on exactly the things that look most checkable. One of us built a generator that took the name of a bioinformatics tool and produced a review of it, and the instruction that ended up in the prompt was “don’t output anyone’s name or don’t output any institution, basically no facts”, because it got names and institutions wrong every time and got them wrong in beautifully finished sentences — “so confidently incorrect that it is worrying”. The working fix for a machine that invents facts was to forbid it from stating any.

So the defence was reading, and everyone arrived at it separately. Reading is cheaper than writing. “You do have to read through it and check it but it’s a lot easier to scan through a file and check something than it is to go and write it from scratch.”

Everyone arrived at the same worry too, which is that the argument only protects people who already know how to program. Hand it to a first-year student who cannot tell a plausible join from a correct one, and what you have given them is not help. It is a very fast way to be wrong.

It is a good argument. It held for exactly as long as the window stayed small, for a reason nobody stated because it was invisible while it was true.

Four thousand tokens

Reading worked in 2023 because the output was small, and the output was small because the input had to be.

A model reads and writes in tokens. A token is not quite a word — punctuation and word fragments get their own — and in English a token averages out at about three-quarters of one. The number that mattered was the context window: one budget shared by your instructions, your code, whatever examples you pasted in, and the reply. In mid-2023 the budget we were working to was about four thousand tokens. Three thousand words, once, for the entire conversation. Whatever did not fit did not exist.

Everything built that year was built around getting under that number.

The literature was starting to trickle in. The study we sat down with1 had put 184 exercises from an introductory bioinformatics course to ChatGPT and marked them with the same automated test suite you would use on a student. It solved 139 first time and 179 of the 184 with further prompting, which was the headline.

The detail underneath was better. The instructor’s own reference answers were never longer than about thirty lines, which is nobody’s idea of a hard problem. Push past that and the behaviour changes. One of us watched it lose the thread on a run of simple percentages and start emitting the same ten lines over and over, which looked at the time like a model that had forgotten what it had already written.

Every tool built on top has the same shape once you know to look. Kellan Vogt* wrote one — write-the, a package that would take a Python project and write the docstrings and the tests for all of it from the command line. By his account most of the work was handling what came back, pulling the docstrings out of the reply and putting them in the right place, and much of the rest was plumbing around the token budget.

It sent one function at a time, in parallel, because docstrings belong on functions, and a function is also a unit that fits. Files too big for the window were split into chunks. Next on the roadmap was a vector database, so the model could at least be told about the parts of the code base it was never going to be shown.

Prompt engineering in that world was a humbler craft than the name suggests. It meant writing out one worked example of a function that adds two numbers, showing the exact shape of the answer you wanted back, and then stopping mid-sentence so the model would carry on rather than start again.

The replies came back in YAML rather than JSON, for the most honest reason anyone has offered — “I just thought curly braces were hard for some reason for LMs.” Validation was proportionate to the stakes. Does it parse as YAML? Then in it goes. The real check came later, when Kellan read the result before sending it upstream as pull requests.

Nobody named the coincidence at the time. Four thousand tokens is a writing budget, but it is also a reading budget, and the two happened to be about the same size. Thirty lines is what a person will genuinely check on a Tuesday afternoon.

For as long as that lasted, the window and the attention span were matched, and every confident thing anyone said about verification was resting on it. It was also, briefly, wonderful. We were “rapidly running out of excuses for why we don’t have documentation.”

A digression about Morrowind

The way to find out whether one of these things is lying is to ask it something you already know cold. So one of us asked it for directions from Caldera to Ebonheart in Morrowind, a twenty-year-old video game, chosen precisely because it is obscure enough to be a fair test. Back came a calm, well-organised route. Go to the Mages Guild. Take a silt strider. Both of those exist. Neither goes where it said it goes. “It sounds very convincing, but it’s wrong.”

The test only works if you have played the game. On any subject where you are a tourist — which is most subjects, for all of us, most of the time — the wrong answer and the right one arrive in identical prose, with the same steady confidence, and there is nothing in the writing to tell them apart. Bad output does not look bad the way bad data does. There is no quality string here, filling up with punctuation to warn you.

But that is also why the small window mattered more than it looked. Thirty lines of Python you can audit against things you already know. Three thousand you audit by running them, which tells you the code does something, not that it does the right thing.

Where it stands

Almost every prediction in these conversations was about capability, and capability went roughly as forecast. The ones that went wrong were about people.

The window went first. In July 2023, half joking, one of us asked whether “if the prompt size just gets bigger and bigger and bigger, maybe we might just actually get better and better code”. At that point thirty-two thousand tokens was something only some people had access to, and a hundred thousand was something people had managed in a research setting.

Four months later the tool a guest demonstrated for us still defaulted to GPT-3.5, with a flag for GPT-4 and an open question about whether the sixteen-thousand-token model had reached the right endpoint yet. Context windows now run to hundreds of thousands of tokens and a million is available. The chunking, the one-function-at-a-time parallelism, the vector store standing in for memory — an entire generation of scaffolding built around a wall that has since been moved several hundred metres back.

The prediction was that the code would get better, and by early 2026 it had, though the bigger window was only part of the reason. Nabil asked Claude to compile Mash into WebAssembly so it would run in a browser, the kind of job that had always ended with the model stuck in a loop of compile, error, change something, compile again. This time he gave it a test as well as a task: ten Salmonella genomes through ordinary Mash, the same ten through the browser build, and the two distance matrices compared. Then he went to make dinner.

An hour later it reported that they matched. “Wait a minute, that wasn’t supposed to work.”

The check Nabil set was not reading the code but comparing two tables of numbers against a tool we already trusted, and the machine ran it itself. That is a better answer than eyeballing thirty lines, and a narrower one than it sounds.

Then hiring. In July 2023 the reaction was immediate — “I’m gonna have to change the way that I hire people” — because take-home technical tests were finished, and the replacement was going to be sitting beside a candidate while they solved something unaided. What you needed to know was whether they could catch the model being wrong. Reading code was the skill being tested for.

The confidence rested on a set of tells, and even at the time we wondered how long they would last. A student at the bottom of the class producing something excellent overnight. American spellings. Text left in that still announced it had been written by an AI. Indentation that changed between files in one project, tabs here and four spaces there, because it had been pasted in from different answers. Essays that opened a paragraph with “in conclusion”, and “no one writes in conclusion, it’s boring”. None of those survived a model that could be told to stop doing them.

By early 2026 the position had turned completely over. “Anyone who says oh I want to learn and understand exactly how this code works and go through line by line and whatever, that’s an immediate no.” Same podcast, same we, three years apart. What was the test in 2023 is the disqualification now, and the person saying it reports building whole projects without touching a line of code.

Perl got a shorter reprieve than anyone expected. Five of the 184 exercises were never solved at all, and one of them was a regular expressions problem — find the words in a passage that start with a vowel — which it failed because it could not reason about the spacing and punctuation in the input. That was enough, in July, for the consolation that “so for all of the Perl people out there, there’s hope for you yet, you won’t get replaced by ChatGPT” — yet.

By November Kellan was demonstrating a sub-command whose only purpose was converting a file from one language into another, with Perl into Python as the worked example, and apologised for probably having lost the two Perl programmers who listen to us. The regex weakness did not survive the year either.

The authorship question was settled without an argument. An editorial Andrew wrote on the ethics of AI in microbial genomics2 — human introduction, machine body — went into the submission system with ChatGPT listed as last author. The name got almost all the way to publication and then died in copy editing.

The argument against was that nobody credits the typewriter, though Andrew wasn’t sure it held for something generative. The best suggestion came from Jimmy Yoon*, a friend of the show, who thought the conflicts of interest section should then have read “one of the authors is an AI and may have conflicts of interest regarding the ethical use of AI”.

What we did not say on air, because it had not filtered through, is that the question was already closed. The Committee on Publication Ethics, the International Committee of Medical Journal Editors and the major journals had all ruled during that same year that a model cannot be an author, on the grounds that authorship means accountability and a model cannot be held to account. The position we were still turning over in public had become policy, and it reached us as a silent deletion rather than as a reply.

The forecast that has aged strangest is the one about two models talking to each other: a paper written by a machine and reviewed by a machine, with the humans watching from the side. In early 2026 Andrew ran the first half as an experiment. Two hundred fields of clinical metadata crossed against genome sequencing data, five thousand combinations tried, which “produced 40 papers of significance” — statistics, figures, references, and by his own reckoning good enough to pass in a journal.

None were submitted; it had been a thought experiment. The reason that matters here came from another panellist at the Bristol session, who guessed that nobody could have checked them well enough to stand behind every page. That is exactly the safeguard everybody nominated in 2023. It works at the rate one person reads.

Set against all of that, the most modest request anybody made in three years of these conversations was aimed at a sequencing manufacturer, and offered free: “get the machine to auto-correct my sample sheet when it’s wrong” rather than only announcing that it is wrong, and fix “the number of commas in the sample sheet” while it is in there. It remains the least likely thing on the list to arrive.

One disagreement should not be tidied away, because both halves came true. One of us expected large language models to improve basecalling and read correction, on the grounds that a genome is a language with rules and predicting the next thing is what these models do. A minute earlier he had joked about a machine that takes your 10x of sequencing and generously invents the other 1000x, and another of us wanted none of it. “I really hope they don’t go that way this starts making up data.”

Basecallers did get transformer models underneath them and accuracy did keep climbing. Generative output that looks exactly like data is now a standing worry rather than a punchline. Both of them were right, which is the least useful way for a disagreement to end.

Then the failure we caught early and mis-scoped. The complaint in late 2023 was that the model kept suggesting deprecated pandas calls, because deprecated pandas calls are what the internet is mostly made of. It was giving us the consensus, and “just because it’s a consensus doesn’t mean it’s correct.”

A newer model may well have caught up with that particular call. The shape of it has not. Whatever changed most recently is what the model has read least about, so it advises you with a slightly stale average, in the same confident voice it uses for everything else.

And the reading problem has quietly got worse in a way none of these episodes saw coming. The 2023 rule was: check the output. By 2026 you cannot altogether trust the input. An AI in the Biosciences panel in Bristol, chaired by Liam Bhatia*, borrowed Simon Willison’s name for it, the lethal trifecta: point an agent at data you did not write, give it access to private data and to the internet, and there was, the panel said, no reliable way yet to stop the private data reaching the internet, because the data you asked it to read can contain instructions.

The document being analysed is also a prompt.

The best formulation any of us has managed is still the oldest one, restated by that same panel as a room full of interns. Quick and capable, and not people whose work you would pass on without reading it. Nobody has ever disagreed with that part.

Which lands this chapter back on the oldest argument in the book. The reason not to write a new tool was never that the idea was bad; it was that somebody has to be there when it breaks. A machine that writes the tool in four seconds has not changed that. It has only made the writing cheap enough that nobody stops to ask who that somebody is.

In June 2023 a model deleted Lee’s name from the top of a short Perl script and we spotted it within seconds, because the script was short enough to take in at a glance. Later that year a copy editor deleted a model’s name from the authors of Andrew’s editorial, and nobody argued about that either. Both edits were caught by somebody reading carefully, in the last year when reading carefully was cheap.

Even on the day, reading was as far as it went.

“I’m not going to know until I run it. I’m probably not going to run it, but it looks right.”

Notes

  1. Piccolo SR, Denny P, Luxton-Reilly A, Payne SH, Ridge PG. (2023) Evaluating a large language model’s ability to solve programming exercises from an introductory bioinformatics course. PLOS Computational Biology 19(9):e1011511. doi:10.1371/journal.pcbi.1011511
  2. Page AJ, Tumelty NM, Sheppard SK. (2023) Navigating the AI frontier: ethical considerations and best practices in microbial genomics research. Microbial Genomics 9(6):mgen001049. doi:10.1099/mgen.0.001049

What Happens When the Money Stops

Outside a hackathon in Bethesda, Maryland, in the autumn of 2024, Tristin Lindemann* was back in front of the same microphone that had caught him five years earlier in Norwich, on low chairs round a coffee table, as the first guest this podcast ever had. A pandemic had happened in the gap, and his assessment of what it had left behind was that everybody now had a sequencer, which put the field “back where we started, but 10 times worse”.

Then he mentioned some news.

“I finally actually got a grant. My first grant ever.”

He knew how that sounded and said so. Prokka had been sitting in almost everybody’s annotation step for a decade by then. Snippy was doing the variant calling in a great many routine pipelines, in laboratories that had never emailed him and never would. Twenty years of software other people’s careers were built on, and this was his first grant — from the Chan Zuckerberg Initiative, paying for the maintenance of things that already existed.

Everybody who has done this job for any length of time has a version of that story, and the versions are all the same shape. Something gets built on a project. The project has an end date, which was decided before anybody knew what would be built or whether it would work. The end date arrives, the person who understands the thing goes somewhere else, and the thing carries on being used by people who have no idea that nobody is behind it any more. Elsewhere in this book that observation is a claim about software lifespans and about why maintenance cannot be funded. Here it is a claim about a date: the date is set first, it is set by an administrative calendar, and everything downstream of it — how long tools live, what shape careers are, who ends up doing the unpaid part — is an adaptation to a deadline that has nothing to do with the work.

The box that did not exist

The interesting part of the Bethesda conversation is not the money. It is what happened when he sat down to write the application.

Research grants had never fitted, and the reason he gives is precise: “they don’t sort of care about the tools, they’re more about the questions you’re answering”. That is not a complaint about taste. A research application is a form, and every box on it asks about a question you intend to answer. A person whose entire professional output is the machinery other people use to answer their own questions has nothing to put in any of the boxes. He had been trying for years and had assumed the difficulty was his.

This form asked what the competing tools were and what state they were in. It asked what was going wrong with his software, and why it needed maintaining, and what he would do if somebody paid for it. He had had those answers to hand for a decade.

“Is this how it feels for academics?”

That line locates the problem somewhere more useful than a general shortage of cash. A funding call is a list of the things somebody has decided are worth describing, and for most of this field’s history there was no box on any form in science that said what is going wrong with your software. The answer every maintainer could have written in an afternoon went unwritten because nobody had ever asked for it. By his account the scheme he applied to was “the first granting program really in the world that’s explicitly supported software”, which, if it is even roughly right, dates the invention of the question to about five years ago.

The trick the field used instead — keeping a tool alive by presenting it as a new one, a version two, the same method moved to a different organism — is elsewhere in this book, and it worked, and it is not a funding model. It is a way of describing maintenance to a form that will not accept the word.

Hours, which are the thing you cannot buy

There were exceptions before that, and they were narrow. The Biotechnology and Biological Sciences Research Council runs a call each year that is “basically not for the new shiny thing but it’s for keeping certain software applications going and certain databases going that have become critical to the industry”. One call a year, heavily oversubscribed, because every group whose grant is running out throws an application at it, and the pot is small in a way that has nothing to do with how much of the country’s science is sitting on top of the tools in question.

And then there is what the money is allowed to touch when you get it. Research funding tends to be earmarked, and the earmarks favour things with serial numbers.

“It’s actually the staff, it’s work hours that you really need to keep something going.”

That sentence is the whole economics of this in one line. You can often find money for a sequencer, a storage array, a rack of compute. You can very rarely find money for a person to spend Thursday afternoon working out why a dependency update broke the parser. Security updates, a library that changed its interface, a language version that stopped being supported — none of that adds a feature, all of it takes hours, and hours are the one thing the budget line will not convert into.

The United States solved the conversion problem by inventing a different kind of employment. A federal or state agency with money allotted to a project, and without the standing workforce to do the work, can contract it out — which is how Nathan Brodsky*, one of the four people who founded the state public health bioinformatics group, came to leave and start a company doing precisely that, and how Ryan Bautista* and Travis Hudak* came to join him four months apart, one out of a state laboratory and one off a federal team, taking on exactly that kind of work for public health agencies. Ryan’s description of the model is that there is a specialised need the government workforce does not have the bandwidth to cover, and that the two of them are, more or less, “hired guns for specialised projects”. Our own summary, from this side of the microphone, was “the Uber of bioinformatics”.

“But with less murder.”

The joke is good and the structure underneath it is not a joke. A contract turns a fixed sum into a fixed-term person, which is the conversion the earmarks forbid — and it does it by making the person, rather than the post, the thing that ends. Everyone involved knows the date. It is in the paperwork before the work starts, and it has nothing whatever to do with when the problem will be finished.

Nobody meters this

Whichever form you are filling in, the last box is the same one, and it is the hardest. Show that this is used.

For a tool that anyone may copy, there is no honest way to answer. Package repositories will give you a badge with an install count on it. GitHub will show clones and stars. A paper has citations, if anybody remembered to cite it, and if the thing they cited was your tool rather than the pipeline it was buried in. But open source has “no control, no centralised control, and so you can’t actually get an exact number always” — anyone may mirror it, vendor it, fork it, copy the one file that matters into their own repository, and none of that reports home.

The failure has a direction, which is the part worth understanding. Every count you can produce is a count of somebody acquiring the software. None of them is a count of anybody running it. And the two diverge fastest exactly where the tool is most load-bearing: once your program is inside a container image, inside a workflow, inside somebody’s validated standard operating procedure, a year of daily production use across a continent registers as one pull of somebody else’s image and no downloads at all. The more thoroughly a tool has been absorbed into the infrastructure, the more completely it disappears from the only instruments anybody has for seeing it. The successful case and the dead case produce the same number.

So people fall back on things that are not measurements. Scrape the web logs, if it is a service and not a program. Run a survey. Start a mailing list — “almost like a boots on the ground technique”, and, when you look at it, a request that users identify themselves voluntarily so that you can prove to somebody else that they exist. Ask around. Read the papers and see whether anybody says what they ran.

Or notice one person.

“One of the people behind BioPerl started watching one of my projects and I was like, oh, that’s a really important person to me.”

Not a number, and not evidence, and yet the only signal in the entire list that anybody actually felt. Nobody has built a meter for that either.

What stops, and what merely rots

When the money goes, a program does not vanish. It sits in a repository, working exactly as well as it did, until the world moves far enough that it will not install, and that can take years. When Nabil went through the numbers from Seebot in 2026 — a code-checking tool, which pulls down bioinformatics code bases and measures them — most tools in our space turned out to get a commit about once a year, and some popular ones we still use had gone about a thousand days without one. They work. A quiet repository is not the same thing as a dead tool.

A service is different. A service needs a machine, a certificate, a domain renewal and — the expensive part — a person to look at what comes in.

The typing databases are the clearest case. Multilocus sequence typing works because the allele numbers mean the same thing in every laboratory that uses them, and they mean the same thing because a human being reviews the sequences going in and decides what is new. Those databases “jump from grant to grant”, and between the grants somebody does the reviewing anyway, which is why the verdict on the arrangement is that “It’s essentially a service given for free.”

That is the failure mode nobody plans for, because it does not present as a failure. The service stays up. It answers. It just quietly stops being curated, and nothing on the screen says so. A tool that dies makes an error message. A reference that stops being maintained makes nothing at all.

Where it stands

The pandemic money did what pandemic money does. In the same Bethesda conversation, Tristin described the state of his own laboratory in Melbourne — one of the more capable national genomics operations anywhere, on the strength of a large injection of cash that had since gone. Budget cuts. Bioinformatics teams smaller than they had been. People who had left public health for industry. And, as a direct consequence, less software coming out of the place than used to. The consortium version has the same arithmetic in a bigger font, and what happened at the end of the British one is elsewhere in this book: stood up in a fortnight, closed on the date the funding said.

What has changed since is that the people building the next thing are asking about it out loud and early. Pathoplexus, the open sequence database Hannah Coxwell* and Alfie Marlowe* were describing in 2025, has a funding page, takes donations, and has found institutional support surprisingly hard to come by. Their appeal ended with the note that if “you’re in touch with a rich philanthropist”, they would like a word. Hannah’s diagnosis of why is worth having: agencies are investing enormous sums in generating more sequence, in more places, from more pathogens, and what becomes of the data once it comes off the machine is left to whoever picks it up — by everybody in the chain, each of them reasonably confident it will not have to be them.

Set against all of that, the arrangement that has actually kept a lot of software alive in this field is not a scheme at all. In 2019, five years before the grant, asked how he kept up with maintaining tools there was no payment for, Tristin gave the honest answer: “I’m very lucky that my employer sees my role in maintaining this software as an important part of my job.” He was quick to add that plenty of his peers had no such arrangement. One manager, deciding that it counts. No line item, no call, no form. It is not a funding model, it is a piece of good luck, and a great deal of what this field runs on has been resting on it for twenty years.

Which is roughly the arrangement that produced the material in this book. Years of conversations recorded in spare rooms and basements and hotel corridors, hosted for a few pounds a year on somebody’s personal account.

“There is zero money — dear advertisement — zero funding agency for this, but if someone wants to buy us a coffee they can.”

Teaching the Thing Nobody Was Taught

Two people came over from Denmark in 2010, or 2011, nobody is sure which, to teach a room at the Centers for Disease Control and Prevention full of people who had never seen a command line how to use one. They arrived with a virtual machine carrying their own toolkit already installed on it, and a list of commands in order. “They just straight up told you exactly what to do and you got publication ready figures”, and people who had come in knowing nothing went away holding something they could put in a paper. The command line was still hard for them. It worked anyway, and the verdict about a decade later was “I thought that was just visionary”.

Nothing about it was a course in bioinformatics. What Harold Oakden* and his colleague had understood, and what most people running training had not, is that a room like that has capacity for about one new idea, and everything standing in front of that idea has to be lifted out and carried by somebody else — the install, the operating system, the data, the order the commands go in. Every course that has worked since has been a version of taking things out, and every argument about training in this field is an argument about which things, conducted by people who were never trained in it themselves and are teaching from a memory of their own confusion.

Five days and a folder

The advanced courses at Hinxton ran to a week, and the week had a shape. Day one opened with an hour or two on the command line, then a visual tool or two — Artemis, usually — and then, having established that a directory contains files and a file has a name, the syllabus went “then it’s like slap bang into things like the inner workings of mappers and assemblers and RNA-seq”. You went home with a virtual machine and a folder of material, both more generous than they sound and neither of which you opened again. What you did not go home as was an expert, because that is “five, 10 years of work of a university education full-time”, and no arrangement of five days closes a gap of that size.

The people arriving know this and want it not to be true, which is a fair way to feel about a fortnight of your funding. The recurring visitor is a clinician who is neither a biologist nor a computer scientist, usually formidably clever, who wants a result somewhere past the end of a PhD and has budgeted a week for the prerequisites. “They want to know everything as quick as possible with the minimum of effort”, and that is not laziness, it is the correct response to a training landscape that has never once said out loud what the price is. The phrase reached for on air for the person on the receiving end was “training that untrained monkey to work as quickly as possible with these really advanced PhD level topics”. That is our phrase, not theirs, and it is not a compliment to us either.

So the week gets designed around an exercise, and the exercise is where it comes apart.

The example nobody can pick

You want the exercise complicated enough to be real and short enough to finish. Those are the same dial and they point in opposite directions. “You can’t just say okay go and assemble your plasmodium falciparum and then expect them to get an answer in five minutes”, so the example shrinks, and a plasmid assembly that lands in seconds has taught the command but not the biology. Both ends of that have been tried, in public, more than once.

“Too simple too complicated and both of them are a disaster.”

There is obviously a middle. The problem with the middle is that it is not in the same place for two people in the same room. As Callum Devlin* put it, the right size and the right amount left abstracted away are “going to be different for different people in the class depending on their learning objectives and their background”. His example is the PhD student who turns up with a hard drive of their own bacterial reads and has brought the exercise with them — “they have an easy complex example that they can kind of build in themselves” — so every command in the week lands somewhere. A master’s student attending because the field is going somewhere and “they hear it’s a up-and-coming field with podcasts coming out all the time” is being given an arbitrary example about an organism they have no stake in, and nothing sticks to it. Two people, one room, one dataset, and only one of them is having the course you designed.

The lever that helps is not the size of the exercise but the door you come in through. A course on antibiotic resistance run with clinicians, pharmacists, wet-lab scientists, industry and a couple of bioinformaticians in it worked, and worked because it led with the domain rather than the methods — the science first, the drug classes and the mechanisms, and the analysis arriving as something that happens to them. Callum, who taught on it, is clear about why the other order would have failed: lead with algorithms in front of a spread like that and “the unbalance in different backgrounds was so much greater”. Biology is the only thing everybody in a mixed room has already agreed to care about, which makes it the only available floor.

What a program will tell you about itself

Since a week cannot hand over the material, Callum’s fallback is to hand over the ability to go and get it: the terms, the man pages, the help flags, where the documentation lives. Nobody disputes this; it is on the first slide of every course any of us have run. What is worth saying out loud is his warning underneath it, which is that when that part does not land, “there is relatively limited long-term retention for those intensive courses”. The folder goes on a shelf and the week evaporates.

Which would be fine if the looking-up worked. It half works. A beginner is taught to type the tool’s name followed by -h, and that is a convention rather than a rule: in some tools -h is a real option meaning something else entirely, and you have just asked for a header line or a human-readable size. “So it’s not universal, and the same with dash dash help”. The next fallback is running the thing with no arguments, which usually prints a usage message — and usually is the whole problem, because it is a good habit only as long as the program does nothing destructive when handed nothing, which you establish by trying it. Man pages exist for the tools that came with the operating system and mostly not for the forty you installed last week. What we tell people to do when they are stuck is the one part of the system nobody standardised.

Underneath that sits a smaller failure that eats afternoons. A beginner working from notes will “blindly copy and paste or blindly type stuff in without really understanding what they’re doing”, and the thing they do not know is that a dash is load-bearing: “it’s not like something in copy and paste from a Word document where it can be a different character”. Paste a command out of a document that has helpfully converted a hyphen into an en dash and the shell does not say the dash is wrong. It says the file does not exist, or prints the usage message. An experienced person deletes the dash and retypes it without noticing they have done anything. A beginner has been handed an error that is true, unhelpful and unconnected to the mistake, and no amount of being told to read the documentation gets them out of it.

Which is why the teaching move that matters most is also the one that looks least like teaching, and it costs time nobody in a five-day course has:

“Not giving the answer when the question is asked.”

It is Callum’s rule for practicals. Somebody puts their hand up, and instead of the flag he shows them where the flag is written down, what the default is, and how he would have looked it up. It is slow, and it is the part of the week most likely to still be there in March.

Which half of the room is behind

The uneven room has a longer version in graduate school, where the intake is half biology and half computer science and the professors have to find a floor for both. It is predictable for about two semesters and then it reverses. “People with computer science backgrounds excelled super quickly at the beginning of the program. It was almost demoralising from my point of view”, from the side that came in through biology — and then the sprint stops, because having the methods is not the same as knowing what they are for, and that half is the harder one. “The biology is really hard to catch up on because it’s a bunch of memorisation”, which is less a slight on biology than a description of how much of it is exceptions with no rule above them to derive them from.

The reverse is the part nobody expects, and it comes from Callum, who taught himself enough computer science to get hired into a computer science department and arrived expecting the students there to be at home on the command line.

“I’ve been surprised at how poor a lot of cs students have been at just command line file munging.”

Not the algorithms. The plumbing: piping one thing into another, moving a file, getting data that arrived in the wrong shape into a program that wants it in the right one. A degree built on theory and clean examples never has a dirty file in it, so the thing everybody assumes the computer scientists came pre-loaded with is the thing nobody taught them either. The two halves of the room are symmetrical: each is behind, in a place the other cannot see, and neither can be told about it by the person at the front, who learned the whole business somewhere else again.

There is a running consolation that says this matters less over time, because the field has matured: nobody needs Linux administration any more, “that’s okay that we’ve abstracted some stuff away”. It survived about four seconds, because virtual machines and containers put it straight back. “Oh okay well okay bad example”. Callum’s version that is not a joke is the postdoc who orders a blade cluster to arrive while they are away teaching a course, and the only person on the corridor with any computational skills goes to the data centre and sets it up.

Where it stands

A panel in January 2021, closing a joint workshop on analysing SARS-CoV-2 data run by the ARTIC network and CLIMB, the Cloud Infrastructure for Microbial Bioinformatics. A clinical microbiologist from Scotland asks the question every course ends on: what should I do next. Rich Haddon*, chairing, hears it the way you hear a thing you have heard three hundred times.

“It’s a very common question at the end of a bioinformatics course, is are you running any more bioinformatics courses, and I always wonder if that just means we’ve done it wrong or it just means no, you’ve done it right.”

The answers run to workflow languages, online problem sets and raising an issue on somebody’s repository, all from Toby Linton*, and the paid Edinburgh courses — and then one nobody expects from a training panel, which is that you should not be the one doing it at all. Hire someone who already knows. A week “might be a nice introduction”, and enough to direct other people’s work if you are managing it, but the doing takes years, because “there’s so many little gotchas and so many little quirks, it takes years and years and years to learn”, and “you might spend six months faffing around when you could just hire someone who can do the same thing in two days”. The chair called it provocative and said he liked it, then added the catch: a year earlier there had been no experts in this particular virus, “and arguably there are still no experts”. He knew exactly where the argument was coming from, and the exchange ended on a joke about the job adverts still to come, asking for three years of experience in it.

That is the part that has dated, and it dated in the least interesting way available. By 2026 the adverts are satisfiable, and the argument has not moved — it has attached itself to whichever technique is now twelve months old. The cycle is older than the pandemic. Computer science filled up in 1999 with people who had heard about the money and “you’ll make a fortune to be a millionaire when you’re 25”, and the first-year attrition was enormous once it turned out to be difficult. In October 2020 Callum could already see the same shape with a new occupant.

“And now we just do the same thing with deep learning.”

Run a deep learning bootcamp, in his example, and you can get somebody running a network on some data “pretty damn quickly but they have no idea what they’re doing necessarily”. Since then the tooling has made the first half of that easier still, and the second half has not needed a word changed.

What has held up is the unglamorous end. A training course in the Gambia gave every student a virtual machine of their own, spun up on CLIMB, preloaded with Conda and a few assemblers, then left them alone with some basic datasets. The Oakden move again, with the environment lifted out so the syllabus would fit. Clodagh Nevin*, remembering a week in Ghana a couple of years earlier with the ARTIC network that ran from sample through sequencing to analysis, gives the flattest possible account of where the time went: “A lot of the time got sucked up actually on very basic bioinformatics.” And the one programme in either conversation with a number attached to it is not a bioinformatics course at all. Micro-research gives people with no research background a mentor and a little seed money and walks them through a whole small project, proposal to publication; by Callum’s account about seventy per cent are still in research five years later, and “most groups end up with a PubMed publication out of it”. It makes no attempt to teach the statistics or the ethics in two weeks. What it hands over is the terminology, the language and the resources “to know how to ask how to do something”.

Which is the same finding arrived at from four directions by people who agree about very little else: the transferable part is knowing how to find out.

Somebody at the back of a course is typing a tool’s name followed by -h, and getting back a usage message, or a header row, or nothing at all, and cannot tell from the screen which of those has happened.

That is the thing we are all teaching around, and nobody has fixed it in thirty years, because fixing it was never anybody’s paper.

Careers Nobody Planned

A university in the American Midwest, 2008, and people keep stopping each other in the corridor holding zip discs.

“My anterior transcriptome is on the zip disc and it’s got all the secrets to my organism and I can’t open it in Excel. What do I do?”

Silas Fenwick* had arrived that year to do embryology. Sea urchins, gene regulatory networks, six or seven years of real bench work behind him, and before that a maths degree and no formal computer science at any point in his life. The plan was to combine good embryology with large-scale gene expression analysis. Then the Illumina Genome Analyzer II became widely available, the corridor filled up with zip discs, and within a couple of years the plan was gone — not abandoned in a crisis, just outvoted by the observation that every biologist on the planet was about to generate more data than they could handle and somebody had better be standing at the other end of it.

Nine people tell a version of that story here. The names matter less than the shape of it: nine doors, and not one of them was a route.

Nine doors

Travis Hudak* and Ryan Bautista* were going to be doctors. Both came into biology the American way, into a department where most of the intake is pre-med, on the strength of two premises — “I like sciences and I want to help people” — and one visible career that satisfied both. One had been aiming at the television doctor from House. They ended up in the same research lab in white coats, and the detail worth keeping is that “our white coats were tie-dyed purple and gold”, in the university colours, handed down afterwards and still in use.

Rich Haddon* was a practising medic who, as Graham Thorne* tells it, got out because the UK government had put more people into junior training than it had jobs for. Graham had five years of funding and could put him on the staff, which brought the PhD fees down to about five hundred pounds, with the option of going back afterwards with the letters. He did not go back and the field of bioinformatics is eternally grateful for that.

Mike Garside*, an applied biology graduate, had to argue his way onto the one undergraduate module with bioinformatics in the title, then took his own PhD title at face value — “genomics was in the title of my PhD but what my supervisors actually meant was AFLP”. He spent it in the wet lab, went down to the bioinformatics group when an RNA-seq chapter turned up and he had no idea what he was doing, and “bothered them until they let me use CLC Bio”. On the first morning of the job that followed he was told to SSH into a server and had to ask what that meant, which is why he thinks the best people to train beginners are the ones who learned it two or three years ago and can still remember the command line being frightening.

Vic Lombardo*, a lab technician, kept handing his data over and then wanting to know what had been done with it, found an old Linux box lying around the division and got the permissions to use it. What settled it was not the Linux box. It was watching an automated liquid handler arrive.

“I am never going to be as good at this as this thing.”

The reasoning is not romantic. He liked the bench, the hands, the sequencers. The machine could pipette all day and all night, “I can only work so many hours in the day before my hand falls off”, and what follows is a practical question about where a person is still worth having: “where can I be more useful than I am in the lab”.

Holly Penrose* ran a Mexican restaurant. “I ran a Mexican restaurant in Melbourne for a year”, after a thesis write-up that produced a breakdown and a dyslexia diagnosis, and she has also run a coffee shop and spent six months teaching chemistry and biology at a girls’ school in Uganda. Farid Tariq* stacked up an undergraduate degree started in 2002, a hospital job, two master’s, a teaching fellowship and a PhD in Birmingham, calls the sequence boring, which it is not, and delivers the only actual warning anybody offers in nine episodes.

“It’s a dangerous job, if you are coming into research you might start liking it, be careful.”

The door did not always open onto something anybody wanted. Tammy Doolan’s* begins with a new bioinformatician told to validate resistance genotypes against phenotypes, who finds the ideal dataset and then finds the phenotypes exist only inside a PDF, and so learns to parse ten thousand samples out of a document. “I must admit it did occur to me that maybe I might have taken up a job I might not actually want in the future.”

What is striking is not that none of it was planned. It is how quickly each of them retrofits a plan afterwards. Mike’s move to another continent within days of his wedding, he and his wife having quit their jobs for a post advertised in a tweet, gets described as part of a long-term strategy about diarrhoeal disease — and then undercut in the same breath, because he knows exactly what he is doing: “This isn’t what I tell fellowship application panels.” Ryan, asked why public health of all things, gives the honest version in four words. “Public health chose me.” A supervisor happened to be talking to a state lab that happened to want its first bioinformatician, and the offer was taken because the post was remote and Ryan was set on travelling in South America first. The fellowship was carried out from Bogotá, and the state laboratory did not hear about it for a while. “I didn’t tell DCLS until a couple of months, a year later.”

The tidiest retrofit belongs to Graham, who won a sequencer. Ion Torrent ran a Europe-wide competition where the prize belonged to you personally rather than to your university, he entered with a proposal about tracking a multidrug-resistant hospital pathogen, and “they swallowed my Blarney and gave me an Ion Torrent”. He collected it in Lausanne and set Rich to working out its data. Then an outbreak happened in Germany, a German group released Ion Torrent data, and Rich analysed it and posted the results on Twitter, where it became a crowdsourced analysis. The New England Journal paper1 was a sprint of its own: the German team needed the analysis redone and the paper rewritten between a Tuesday and the end of Thursday. Told afterwards this becomes chance favouring the prepared mind, and the verdict is “we were in the right place we’d lined up the things” — true, and also a sentence you can only assemble in that order once you know how it ended.

The thing that came out of an accident

The best technical idea in these nine episodes exists because Silas, with no computer science training, had built himself a Python interface for unrelated reasons.

The setting is a single-cell E. coli dataset sequenced absurdly deep, and a car journey home. “Why do we need 400x coverage”, when fifty to a hundred would do the job better. This is not an aesthetic objection. Depth past a point stops helping an assembler and starts hurting it — more to hold in memory, more to reconcile — and “the assemblers often fall over or they pick the wrong thing to focus on”. You can cull reads at random to bring the depth down, which is fine on a clean isolate and a disaster on anything mixed: a random cull takes the same proportion from the organism you sequenced a thousand times over and from the one you barely caught at all.

The idea that arrived on the drive home was to cull by coverage instead: “estimate the coverage of each read and just throw it away once it got high enough”. Textbook version: digital normalisation2 estimates each read’s coverage from the counts of the k-mers inside it, keeps the read if the estimate is below a threshold and discards it if it is above. Honest version: it is a way of not reading the same sentence four hundred times.

The consequence is the part worth having. Because the threshold applies per read rather than per dataset, the abundant organism sheds nearly all its reads and the rare one sheds none, having never been above the threshold in the first place. That is what makes it work on a metagenome rather than only on an isolate, and why the working benchmark was soil, on the principle that “if you can analyse a soil metagenome, any other metagenome is easier”. It also comes with a bill. Once reads have been deleted on the basis of how many of them there were, what survives is not a count any more. Normalise for assembly, not for quantification.

The implementation took about thirty minutes, in the gap before dinner, and the reason it took thirty minutes is the point of this chapter. There was already a clean Python layer sitting on top of the fast code, written years earlier by the same person, who had come to biology through maths and open source and who cared about interfaces for a reason that was social rather than technical. He had once built a comparative sequence analysis tool for his lab, plus a front end, plus a tutorial, then gone back to his own work — and turned up at a meeting eighteen months later to find that eight people had used it, got results and were ready to publish without ever having had to speak to him.

“This is the way to do bioinformatics, I get to write software and that I don’t have to talk to people.”

That is a joke, and also a design philosophy with a body count of good tools behind it. Writing documentation in order to avoid conversation is not a motive anybody puts in a grant application, and it produced the Python interface, and the interface is why the coverage idea took half an hour rather than half a year.

The community had to be built before anybody could join it

For a long time the physical fact of this job in the UK was a room with no windows. Mike spends a while gently correcting the record — not technically a basement, a room at the back with an overhang outside it and a backup generator, hence no natural light — before conceding it entirely: “it felt like a basement and it smelled like a basement”.

The argument worth having is whether you should be in it. Two models, then and now: put every bioinformatician in one central group, or keep a hub and embed spokes among the microbiologists and epidemiologists who use the results. Mike had been a spoke rather than part of the hub, embedded with the gastrointestinal bacteria unit, and comes down firmly for the spokes — an assay is more likely to be useful if the person building it talks daily to the people who will act on it. He then supplies, unprompted, the best evidence against his own position. He once spent “a horrible month of my life writing a Python application” that emitted the XML needed to push ten thousand genomes a year into the NCBI pathogens portal, and found out three months later that a bioinformatician downstairs had just spent a month writing the identical thing. He would still take the spoke. A real cost, honestly stated, and declined.

Scattered far enough, the model stops having a hub at all. Mike’s own career ends up in Malawi as “the only bioinformatician in the entire country”, working over 4G billed at about a dollar a gigabyte, in a city on six to eight hours of scheduled power cuts, running everything remotely because “you definitely don’t want to be downloading all of your BAMs for visual inspection”. What replaces the corridor is Twitter and a Slack group, described without irony as still performing an important function. An American public health fellowship hit the same wall from the other side: in one early cohort there was no bioinformatics mentor to be had at any state laboratory, so the community had to be built before anybody could join it. A few years later a new fellow arrives to find a whole population of practitioners already there, whether or not their own institution employs a single one.

All of which is scaffolding, built at speed, by people who had none.

Everybody’s advice is a self-portrait

At the end of the wet-to-dry conversation somebody asks the standard question about what you wish you had known, and the answers arrive so fast they overlap. The hosts start guessing before the guest can speak: “learn Snakemake”, learn Galaxy, learn how to program. The guest, who came out of the wet lab and had bootstrapped his workflows together in Python, says he wishes he had “really engaged with a workflow language”. One host, who came in from computer science, says the opposite — “interacting as much as possible with biologists is probably the key for me” — and a second, from public health, says talk to the epidemiologists too, then adds the concession that gives it away: “coming in with biology knowledge and public health knowledge is just so priceless”.

Nobody in that exchange is describing a route. Everybody is describing the wall they personally hit. Mike, from the bench, prescribes programming, the programmer prescribes biology, the one who learned it late recommends learning it early, and Vic tells you to start on a cloud machine, because a laptop battery once died on him mid-script, and to carry redundancy in everything — two internet connections; headphones, earbuds and a microphone. It is the most consistent finding across nine episodes, and in that exchange it is noticed only in passing, as the same advice given in reverse.

The same inversion is what makes the postdoc argument unresolvable, and worth reading. Lucy Farrant*, Holly and Tariq sit in front of a room of first-year PhD students, Andrew chairing, and are asked why on earth anybody would do this, given the pay. They answer, in order: because the problem is unsolved and nobody else has solved it; because “I think happiness is more than wage”; because a contract that ends is also a contract you can leave, which is how one of them got to Uganda and Melbourne. There is a straight look at the money, from someone whose family had not been to university — “wow I earn more than my parents” — and one at the school friends who are now managing directors: “you’re aging as I look at you and you’re melting like a candle”. There is blue hair, which would not survive a corporate dress code, or a high street one.

“I’m not even allowed to work at Primark, mate.”

And then the host, who has done both, closes the session by telling a room of PhD students that if what they are after is money they should go into bioinformatics and data science rather than the lab, because the earning potential is elsewhere. Nothing reconciles those two positions. They were said fifteen minutes apart in the same room, both are correct, and that is why nobody can hand anybody else a route.

Where it stands

The people asking the questions have moved as much as the people answering them. Two of the three left the institute they were at when most of these were recorded: Nabil-Fareed to a pathogen surveillance centre in Oxford, Andrew to a contracting company in the United States and then a start-up doing colorectal cancer diagnostics — “so not even microbes” — a sentence with fifteen years of microbial genomics behind it and no distress in it at all. Silas had already given the reason, about his own arrival in a veterinary school: “everything has DNA. So if you work with DNA sequence, you can fit in pretty much anywhere.” He was hired there, on his own account, because the department “wanted somebody that they could have coffee with whenever they wanted to talk about data science or sequence analysis”, which is not a job description anybody writes down.

Two things have dated. The first is the substrate. In 2020 Twitter is how you find a job in Vietnam, how a crowdsourced outbreak analysis in 2011 reached enough people to become a paper, and how the only bioinformatician in a country stays in a conversation at all. By 2025 the line in the same podcast is that we cannot even say Twitter any more, and the honest position is that nothing which followed it does that job. Two of these careers were routed through a website that no longer exists in the form that routed them.

The second is money. In May 2021 Ryan, with money arriving on a scale he doubted he would see again in his lifetime, calls it “our turn at bat”, remembers when the whole field could not fill a ballroom at its own conference — “we didn’t even really fill up the main ballroom” — and argues that the responsible use of a wave like that is to build for the diseases not paying for it. Said at the top of the wave, it reads differently from underneath.

One prediction the book will not pretend to have followed up. In January 2023 Vic describes giving up an apartment in Atlanta, putting what is left into storage and working from Airbnbs indefinitely, with the apartment going the following summer. He had already worked from Mozambique, a bedroom in New York, a pizza place and a coffee shop, and had replaced his to-do list with his calendar: “I don’t need to worry about what I’m doing that day because my calendar will tell me.” Whether the apartment went, we do not know.

The door problem itself has changed. There are fellowships now, and taught master’s programmes, and job titles that existed nowhere in 2010. Somebody starting today can be pointed at a route, which is good, and which turns these nine stories into period pieces rather than instructions.

Every one of those episodes opens with the same line, read out before anybody has said anything. “There is no manual, and it’s assumed you’ll pick it up.”

Nine people picked it up. In a windowless room that smelled like a basement, in a corridor full of zip discs, off a tweet, off a liquid handler that was better at pipetting than a human hand, and in one case by way of a Mexican restaurant in Melbourne.

Notes

  1. Rohde H, Qin J, Cui Y, et al. (2011) Open-source genomic analysis of Shiga-toxin-producing E. coli O104:H4. New England Journal of Medicine 365(8):718-24. doi:10.1056/NEJMoa1107643
  2. Brown CT, Howe A, Zhang Q, Pyrkosz AB, Brom TH. (2012) A reference-free algorithm for computational normalization of shotgun sequencing data. arXiv:1203.4802

Genomics at the Cinema

John Hammond’s walking stick has a ball of amber on the top of it, and inside the amber is a mosquito, and that mosquito is the whole film. Blood in the gut, dinosaur DNA in the blood, park on the island, everybody eaten by the third act.

The trivia we turned up with was that “the mosquitoes are clearly male in the amber when they claim that all the animals are female”. Male mosquitoes drink nectar. They do not bite, they have never taken a blood meal in their lives, and you can tell one from across a room by the antennae. So the object the entire park rests on could not have supplied a single base of it, in a park whose one stated safety guarantee is that every animal on the island is female.

That is what got us. Not the sixty-five-million-year-old DNA. The prop.

The DNA was, if anything, the part the film handled well. Getting genetic material out of an insect trapped in amber was not a stupid idea in 1993 — there were publications, the premise had a real precedent, and a paper claiming intact DNA from an insect in amber1 came out the day before the film opened. Later attempts to repeat that kind of result have failed, which is not the film’s fault. What the film also gets right is the damage. It says the code has holes in it, which is the correct shape of the problem: ancient DNA is “really chopped up”, and the gaps are not a plot device, they are the entire experience of working with old material. The film then fills the gaps with frog, which is where it stops being a documentary and starts being a monster movie, but the shape of the difficulty is honest.

So the impossible premise got a pass and the insect’s sex did not, and that is the shape of every one of these conversations. Twice now we have sat down at Christmas to watch a film about our own job, and twice we have not managed to watch one.

Nobody can just look at a lab

Before the tour even reaches the laboratory there is an objection only this room would file. The park rebuilt animals sixty-five million years dead and gave them none of their bacteria. Every one of “all the microbes that those dinosaurs should have had”, in the gut and over the skin, has spent the intervening time evolving into something else, so what comes out of the hatchery has a modern microbiome and probably no reasonable defence against a modern pathogen. The film supplies the case itself and then drops it. A triceratops is sick, its dung gets examined, the poisonous plant is ruled out, and “I don’t think we ever find out what’s wrong with it. It’s just sort of sick”, because the tyrannosaur gets loose and the question stops mattering.

What happens once the tour does reach the laboratory is that we stop following the plot and start reading the set.

The freezers have clear glass doors, like the fridge in a petrol station, when the real thing is a blast shield you cannot see through. The benches sit at angles nobody would arrange them at. The surfaces are clean, and worse than clean — nothing is precariously stacked, nothing is in a box with masking tape on it and a handwritten label that stopped meaning anything in about 2004. The robotic arm turning the eggs can only reach the first two eggs. The embryos get carried around at a temperature that would take your fingers off, in a can, by hand. The tubes are labelled velociraptor, which is not a sample ID, which is not how anybody has ever labelled a tube.

And the geneticist, mid-extraction, having drilled into a lump of amber with no eye protection whatsoever, puts on a full face visor to go and look at a computer screen.

“He has PPE for the computer.”

None of this is where the film breaks, obviously. All of it is a set dresser making a shot look good, and the honest defence kept surfacing every time somebody complained — do it properly and “but then there’s no movie”. The point is that we could not not see it. The paused frame with the DNA background on it got examined for whether the bases were real sequence and worth putting through BLAST, which would have been a genuinely good afternoon’s work. The line about three billion instructions sent one of us into NCBI’s taxonomy looking for dinosaurs, and there is an entry, and the excitement lasted about four seconds, because “the only genomes under dinosaur are birds”. Every bird genome, filed under dinosaur, by someone who had thought about it properly.

The high point is Dennis Nedry’s screen. For roughly a second and a half you can see the code that is about to destroy the park, and it is not gibberish. It is real, and it is version control, and what it actually does is check that you have written a commit message and that nobody else has the trunk checked out. The saboteur’s terrifying screenful of code says “make sure you’ve commented your commits”.

Which means somebody on that production went and found a working snippet instead of typing nonsense, and thirty years later we paused the clip to read it, and the finding was that the snippet is fine and whoever wrote it did a good job. This is the only sense in which anybody wins these arguments.

What a tree can and cannot tell you

The other film is Contagion, which is the good one, and which is treated far more roughly precisely because it is the good one. Nobody had a serious complaint about the epidemiology. The rating offered was “zero out of ten totally unrealistic didn’t see anyone buying toilet paper”.

Some of it is not a set at all. The exteriors are the real Centers for Disease Control and Prevention (CDC): the driveway the car comes up, the visitors’ car park, the museum any member of the public can walk into. Lee watched the crew dressing the site on his way to work, arranging the trees and positioning the fake protesters. The rest of us had parked in that car park ourselves, and recognised it on screen. The mistake was Lee’s catch as well. The film’s agency director arrives in the morning and parks in the visitors’ lot. “I’m not gonna divulge where the director parks”, but it is not there.

Then, an hour and seven minutes in, the bioinformatician opens an email, looks at a phylogenetic tree and rushes off, and on the strength of that tree the agency director is told there is a new R0.

This is the one that should not be let go, because it is a mistake with a working analogue. R0 is not a property of a sequence. It is a property of a virus meeting a population: how long you are infectious, how many people you meet while you are, and how the thing gets from you to them. The film had already explained this, in its own dialogue, and explained it well: the number depends on the infectious period, the contact rate and the mode of transmission. A genome tells you about relatedness. It can tell you that two cases are linked, and that a lineage has moved somewhere new. It cannot hand you a contact rate. You get that from people knocking on doors, which the film also depicts, thirty minutes earlier, correctly.

The tree itself does not survive contact either, and this is where it turns properly forensic. A phylogeny of something ripping through a population early in an outbreak looks like a ladder: dense sampling, near-identical tips, long thin trunk with short branches coming off it, because there simply has not been time for much to change. What is on screen is the other thing entirely. It has long branches, it has tidy well-separated clades, it looks like a figure from a textbook about primates. “I would have imagined that their branches would be very very short.” A tree that stable is a tree of an organism that has been sitting around diverging for a long while, which is the opposite of the film’s entire premise.

And the branch lengths cannot be checked anyway, because there is “no scale bar no bootstraps”. Without a scale bar a branch length is decorative. The dialogue says one cluster is highly divergent — that is the whole reason for the scene — and there is no branch on the picture that is divergent from anything, and even if there were you could not tell how divergent, because nobody drew the ruler.

Then somebody zooms in. Because of course we found the second tree too, the one the other scientist carries into the meeting on a laptop, and established that it is the same tree, zoomed into the blue cluster, and that it still does not show what the dialogue says it shows. The tip labels are illegibly long, which was ruled the single most realistic thing in the scene.

The best moment in the whole two hours comes in the middle of this. Deep into the branch lengths, one of us stops to ask another whether he ever had a bowl cut like the actor’s — a proper detour, resolved amicably — and then, with no transition at all: see, there’s no scale bar here. Nobody remarked on it. Nobody has ever remarked on it.

For the record, fairness did break out eventually, in the form of an observation that “we should be happy that there’s a phylogenetic tree in a mainstream movie to begin with”. It arrived after eight minutes of frame-by-frame analysis, and it did not stop the analysis.

We also worked out that the tree is not being displayed in any tree viewer. It is a picture, opened out of an email attachment, and the toolbar visible along the edge is Microsoft Paint. “It’s just Paint because it’s got paste at the left.”

The software that does not exist

Which brings us to the genome dashboard, and to the least defensible and most enjoyable six minutes any of us have spent on a film.

The scene is a screen. It has four panels: an alignment across the top that looks like ClustalW output, a rotating protein structure, a recombination map showing where the bat sequence stops and the pig sequence starts, and a similarity plot against related viruses. It is, by consensus, a good interface. It is more informative than most things any of us have shipped. Somebody involved in that production had clearly spoken to a bioinformatician, and the resulting layout is what you would actually want.

Nobody recognised it. Nobody could place it, nobody could name a tool it might be, and the first finding was simply that “I don’t think it exists”. That should have been the end of it.

Instead the investigation opened. The window furniture puts it on Windows, Vista or 7, which is not realistic for a research tool of that vintage — except that the panels look reminiscent of Jalview, and “Jalview being a Java program, it would run on Windows so maybe that makes sense”, so it is realistic after all, and the objection was withdrawn on the strength of a runtime. From there, the buttons. The buttons have a look. The look belongs to a particular Java widget toolkit, and retrieving the name of it took a couple of minutes and several wrong answers, including a guess at Qt, which is not a Java toolkit at all. Swing. It was Swing. And then, about a fictional program in a fictional pandemic in a film from 2011: “We’ll look it up and find out. Yeah. We don’t want to get that wrong.”

There is an Internet Explorer icon in the taskbar.

What is worth noticing is where this lands. Not on a complaint. The recombination panel, the one showing the crossover between the two source viruses, is a thing none of us had seen in any real tool, and the reaction to it was immediate and slightly wistful: “that might be a fun thing that we should write this tool”. Nobody could find out how the screen was made, and the best guess was that “some poor bugger actually implemented” a convincing facade in Java, in about 2010, for a few seconds of screen time. The only thing anybody wanted to do about it was find them and tell them it was good.

Contagion is now a period piece

Contagion came out in 2011 and was shot the year before, which is a specific and now slightly astonishing moment. The MiSeq had barely arrived; one of us did not have one until 2012. The sequencers on the ground were Genome Analyzer IIs producing 36-base single-end reads, which was enough to do the Haiti cholera genome, and it’s easy to forget now how much anybody got out of reads that short. The film’s bioinformatician wanders in and out of the containment lab, and in 2011, with fewer than ten or twenty dedicated bioinformaticians at the CDC at all, that was probably true to life. It is not true now, and the standing recommendation for any remake was clear: “You should not see the mathematician go in the lab.”

Mathematician. Twenty minutes of discussion about a bioinformatician character, and the slip that came out was mathematician — which is precisely the complaint filed against Jurassic Park, in almost the same words, about somebody in a white boiler suit with cables round his shoulders working in a wet lab. Thirty years, two films, one grievance: the mathematician is in the laboratory and nobody can explain why.

The other thing that has dated is more interesting than a tool version. Contagion was watched as prophecy for most of 2020, and from this side of it the virus holds up — the fictional MEV1 is modelled closely on Nipah, down to the genome size. The institutions are where we split.

One of us had a problem with how competent government is in the film, having since watched real ones take forever to do anything; another liked seeing government step in and help. Decisions get made in one meeting rather than forty, and the squashing makes the official worrying about the biggest shopping weekend of the year look callous, which is unfair: the science can say that something is circulating, and what to do about it is a matter of priorities. Everyone who behaves unethically gets away with it, which is treated in the film as realism and struck us as a slander. And the thing that would now be the first shot of the first act, the run on toilet paper, is nowhere.

The traffic between the film set and the real thing ran the other way as well. When the United Kingdom put five hundred beds into the biggest conference centre in London in the spring of 2020, some of the ventilators came off film sets, where they had been props that happened to be working machines. It was staffed with retired doctors, junior doctors, vets and physiotherapists, because the primary care physicians could not be taken off the front line. It was barely used.

One habit from those days survives on purpose. Long and short were both measured against instruments that have since gone, so one of us will not say short read at all any more, and says sequence read instead — a term declined out of respect for a machine that is now landfill.

We watched both of these voluntarily, at Christmas, for fun. The dinosaurs were all female. The mosquitoes were not. Nobody has fixed it in thirty years, nobody is going to, and we will check again next time.

Notes

  1. Cano RJ, Poinar HN, Pieniazek NJ, Acra A, Poinar GO Jr. (1993) Amplification and sequencing of DNA from a 120-135-million-year-old weevil. Nature 363:536-8. doi:10.1038/363536a0

Thanks

A podcast is three people talking. This book exists because a great many other people agreed to come and talk to them, usually for nothing, often at a conference, occasionally down a line so bad that half of what they said had to be reconstructed afterwards.

Names below are taken from the episode notes, which were written by hand at the time and are the most reliable record of who was in the room. Where somebody appeared several times, the episode numbers are listed together. Affiliations are left out on purpose: people move, and an employer that was right in 2020 may be wrong now, or somebody else’s to approve.


Guests

Mads Albertsen — SARS-CoV-2 sequencing in Denmark (51, 52)

Nabil-Fareed Alikhan — the BLAST Ring Image Generator, EnteroBase, and the Mastodon migration (28, 45, 94)

Frank Ambrosio — nomadic bioinformatics (98)

Muna Anjum — the Nanopore-versus-Illumina panel (18)

Martin Antonio — the Nanopore-versus-Illumina panel (18)

Phil Ashton — from the wet lab to the dry lab, and a career in Malawi (34, 35)

David Baker — the Nanopore-versus-Illumina panel, and CoronaHiT (18, 23)

Kate Baker — antimicrobial resistance, the deep dive (14)

Mark Basham — AI and the biosciences (149)

Dany Beste — Mycobacterium tuberculosis (64, 65)

Kai Blin — identifying variants of concern by Sanger sequencing (55, 56)

Titus Brown — from maths to metagenomics, sourmash, k-mers, taxonomy (117, 119, 120, 121)

João Carriço — ontologies and data sharing (38)

Clint — guest host for the Pathoplexus episodes (144, 145)

Bede Constantinides — Deacon (151, 152)

Natacha Couto — One Health, and the early days of multi-locus sequence typing (101, 102)

Tim Dallman — the hackathon interviews (130)

Henk den Bakker — Sepia, and the director’s cut (74, 88)

Ed Feil — One Health, and the early days of MLST (101, 102)

Heather Felgate — mobile genetic elements, and the postdoc panel (105, 106)

Kelsey Florek — StaPH-B (78, 79)

Emma Griffiths — ontologies, contextual data, GMI (26, 37, 38, 111)

Ozan Gundogdu — the Nanopore-versus-Illumina panel, and Campylobacter (18, 62, 63)

Verity Hill — SARS-CoV-2 phylogenomics (41)

Suzie Hingley-Wilson — chaired the Nanopore-versus-Illumina panel, and Mycobacterium tuberculosis (18, 64, 65)

Emma Hodcroft — Nextstrain, Mastodon, Pathoplexus (86, 87, 94, 144, 145)

Kristy Horan — the frontlines panel (99)

William Hsiao — SARS-CoV-2 surveillance in Canada, and the Global Microbial Identifier conference (53, 54, 112)

Phil Hugenholtz — bacterial taxonomy (69, 70, 71)

Zamin Iqbal — comparative genomics (77)

Tue Jorgensen — identifying variants of concern by Sanger sequencing (55, 56)

Abdoulie Kanteh — invasive non-typhoidal Salmonella (50)

Curtis Kapsak — StaPH-B containers, and a career in public health (57, 58)

Kostas Konstantinidis — ANI and metagenomics (124, 125)

Christine Lee — the Haiti cholera outbreak (128)

Kevin Libuit — StaPH-B containers, and a career in public health (57, 58)

Nick Loman — the Nanopore-versus-Illumina panel, Majora, SARS-CoV-2 phylogenomics, overcoming barriers to data analysis, and chairing the ARTIC protocol session (18, 32, 41, 42, 44)

Jennifer Lu — Kraken (103, 104)

Duncan MacCannell — the PHA4GE contextual data specification (26)

Grant Mackenzie — invasive non-typhoidal Salmonella (50)

Finlay Maguire — training, SARS-CoV-2 surveillance in Canada, the antimicrobial resistance panel, the Global Microbial Identifier, clinical metagenomics, the frontlines panel (33, 53, 54, 76, 99, 111, 140)

David Mahoney — antimicrobial resistance in metagenomes (132)

Leonardo de Oliveira Martins — phylogenetics with the arborists, Bayesian phylogenetics, and Nextstrain (11, 12, 16, 17, 87)

Alison Mather — sequencing Norfolk (31)

Conor Meehan — phylogenetics with the arborists, Bayesian magic, tuberculosis, bacterial taxonomy (11, 12, 16, 17, 64, 65, 67)

Sam Nicholls — Majora, and tracking SARS-CoV-2 nationally (32)

Justin O’Grady — standing up SARS-CoV-2 sequencing, and CoronaHiT (21, 23, 31)

Mark Pallen — the Nanopore-versus-Illumina panel, bacterial taxonomy, a career either side of the millennium, and the antimicrobial resistance panel (18, 69, 70, 71, 76, 82, 84, 85)

Marike Palmer — prokaryote systematics and SeqCode (107, 108)

Danny Park — the Workflow Description Language (89)

Elisa Pedone — AI and the biosciences (149)

Robert Petit — Bactopia (72, 73, 129)

Megan Phillips — tetracycline resistance in MRSA (138)

Radosław Popławski — Majora, and tracking SARS-CoV-2 nationally (32)

Anna Price — SARS-CoV-2 phylogenomics (41)

Leighton Pritchard — bacterial taxonomy and ANI (67)

Nikhita Puthuveetil — assembling reference strains (135)

Joshua Quick — the ARTIC protocol (44)

Natalia Rincon — Kraken (103, 104)

Miguel Rodriguez-R — prokaryote systematics and SeqCode (107, 108)

Will Rowe — overcoming barriers to data analysis (42)

Theo Sanderson — Pathoplexus (144, 145)

Torsten Seemann — writing good software, lost in translation, clinical metagenomics, Snippy, the frontlines panel (6, 95, 99, 131, 140, encore 06)

Abdul Sesay — the Nanopore-versus-Illumina panel (18)

Joel Sevinsky — the Workflow Description Language (89)

Kieren Sharma — co-hosted the AI and biosciences panel (149)

Sam Sheppard — why use genomics in an epidemic (43)

Iain Sutcliffe — bacterial taxonomy (69, 70, 71)

Brooke Talbot — MRSA and public health (134)

James Thomas — AI and the biosciences (149)

Ruth Timme — contextual data specification, GMI (26, 113)

Clement Tsui — the antimicrobial resistance panel (76)

Niamh Tumelty — naming SARS-CoV-2 variants (39)

Anthony Underwood — the antimicrobial resistance panel (76)

Peter van Heusden — SARS-CoV-2 tools and variants of concern (48, 49)

Arnoud van Vliet — the Nanopore-versus-Illumina panel (18)

Cynney Walters — the Haiti cholera outbreak (128)

Emma Waters — mobile genetic elements, and the postdoc panel (105, 106)

Wytamma Wirth — write-the, and code documentation with large language models (114, 115)

Archie Worwui — the Nanopore-versus-Illumina panel (18)

Muhammad Yasir — chaired the mobile genetic elements panel (105)

Erin Young — StaPH-B, and bioinformatics in public health (78, 79, 133)

Sara Zufan — Lassa virus genomics, and a PhD journey (139)


Also

The CLIMB-BIG-DATA and ARTIC network panels (41, 42), the hackathon audiences in Bath and Bethesda, and the first-year PhD students who sat through a frank discussion of whether to do a postdoc and asked the hardest questions in it (106).

Whoever it was at CDC who told four people to stop bothering her and go and talk amongst themselves, thereby founding StaPH-B. She is named on their own About page, and did not know it until somebody told her.

Every listener who wrote in to correct something. Several of those corrections are in this book, and at least one of them is in a chapter that would otherwise have been wrong.

And the contributors we have deliberately left unnamed, because being named could cost them something with their own government. Their work is in here and their names are not. We would rather thank them properly, and cannot, and we are not going to pretend that is a small thing.


Names are as the episode notes give them; where the notes were silent we have said so rather than guessed. Affiliations are deliberately omitted — people move, and an employer named in 2020 may be wrong, unwelcome, or somebody else’s to approve. Any errors in names or attributions are ours.

A Note on Sources

Everything in this book comes from the MicroBinfie podcast: a hundred and fifty-six episodes recorded between 2019 and 2026. The recordings are at soundcloud.com/microbinfie, and the episode notes — which name the guests, link the papers, and frequently explain what an episode was for in a way the audio does not — are at microbinfie.github.io.

If a chapter makes you want to hear the argument rather than read it, the episodes behind each one are listed below. They are worth going back to. A transcript flattens the timing, and the timing is most of why any of this was funny.

Quoted speech is verbatim at the word level. Punctuation inside a quotation is ours, because speech does not come with any, and where a good line turned out to be two people finishing each other’s sentence we have either cut it back to the part one person actually said or shown the join. Nobody has been tidied up, and nobody who turned out to be wrong has been quietly corrected — where a prediction dated badly, the prediction stands and the correction follows it.

The guests, and the colleagues they talk about, appear in these pages under pseudonyms, each marked with an asterisk (*) the first time it appears in a chapter. The three of us do not, and nor does anybody named only for their published work. The words are still the guests’ own, and they are thanked under their own names at the back of the book.

Where the prose points at a specific published study, there is a numbered note at the end of that chapter with a DOI or an equivalent identifier. Tools mentioned in passing are not cited; a reader who wants Prokka can find Prokka.

Some episodes appear in no chapter at all. Roughly twenty of them are a near-continuous weekly record of a genomics emergency as it happened, recorded without hindsight, and they are a different book. The rest is podcast furniture: trailers, birthdays, conference round-ups.

Chapter Episodes
1. Things That Went Wrong 9, 13, 15 and 33
2. Software Not To Write 1, 2, 6, 24, 25, 95 and 151, and the encore re-release of 6
3. What Language Should I Learn 91 and 92, and the encore re-release of 91
4. The Formats Nobody Designed 3, 7 and 10
5. How This Field Got Here 4, 5, 82 and 84
6. Early Days of MLST 102 and 45
7. Assembly, and Other Acts of Faith 19, 20, 22 and 141
8. Sequencing Technologies, and the Arguments About Them 18, 55, 56, 97 and 141
9. Trees, and How to Argue About Them 11, 12, 28, 126 and 127
10. Bayesian Magic 16 and 17
11. Sketches, Hashes and k-mers 29, 103, 104, 120, 121, 122 and 123
12. What Is a Species Anyway 67, 68, 69, 70, 71, 107, 108, 119, 124 and 125
13. Naming Things 39, 60, 111 and 135
14. Pangenomes 36 and 93
15. Mobile Genetic Elements 83 and 105
16. Resistance 14, 76, 77, 132, 134 and 138
17. Bugs With Personalities 50, 62, 63, 64 and 65
18. Outbreaks 128, 133, 139 and 140
19. When It Was Not a Drill 21, 23, 32, 43, 44 and 87
20. The Pipeline Wars 72, 73, 89, 90, 129, 146, 147 and 148
21. Containers, and Other Ways to Stop Suffering 57, 74, 78, 79, 88 and 152
22. Ontologies, and Other Things Nobody Thanks You For 26, 37, 38, 112 and 113
23. Who Gets to Sequence 18, 26, 41, 47 and 150
24. Publishing, Reviewing, Shouting Into the Void 61, 66, 75 and 94
25. Machines That Write Code 109, 110, 114, 115, 116, 149 and 150
26. What Happens When the Money Stops 6, 24, 25, 30, 57, 131, 145 and 153
27. Teaching the Thing Nobody Was Taught 33 and 42
28. Careers Nobody Planned 34, 35, 58, 85, 98, 99, 106, 117 and 143
29. Genomics at the Cinema 118, 136 and 137

Glossary

Definitions as the words are used in this book, which is not always how they are used elsewhere. Several of the entries below are contested in the field; where that is the case the chapter concerned says so at more length.

Accessory genome. The genes present in some members of a species but not all. What an organism is carrying this week, as against what makes it that organism. See pangenome.

Allele. One of the alternative sequences found at a given locus. In multi-locus sequence typing an allele is deliberately reduced to an integer: two sequences are the same allele or they are not, and how different they are is not recorded.

Amplicon. A stretch of DNA copied up by PCR, usually as a deliberate first step so that a small target can be sequenced from a sample containing very little of it.

AMR. Antimicrobial resistance.

ANI. Average nucleotide identity. Take two genomes, find the parts that correspond, report the mean percent identity across those parts. Ninety-five per cent is the conventional species boundary. Which parts correspond is a property of the software.

Antibiogram. The pattern of which antibiotics kill an isolate and which do not. Each drug is reported as susceptible, intermediate or resistant.

ARTIC. A published protocol, and the network around it, for sequencing viral genomes directly from clinical samples using tiled PCR amplicons. It is how most SARS-CoV-2 genomes were produced.

Assembly. The reconstruction of a genome from the short fragments a sequencer produces. Also the resulting sequence. An assembly is a hypothesis, not an observation.

Basecalling. Turning the raw physical signal a sequencer measures — a light intensity, an electrical current — into letters. The step where most of the errors are made and most of the improvements have happened.

BLAST. Basic Local Alignment Search Tool. The program, first published in 1990 and substantially rewritten in 1997, that finds regions of similarity between sequences. Still the default answer to “what is this”.

cgMLST. Core-genome multi-locus sequence typing. The same idea as MLST extended from seven genes to the whole core genome.

CLIMB. Cloud Infrastructure for Microbial Bioinformatics. The UK academic compute platform that carried a large share of the country’s pandemic sequencing.

Clade. A group on a phylogenetic tree consisting of an ancestor and all of its descendants.

Contig. A contiguous stretch of assembled sequence. An assembly is a set of contigs, and the breaks between them are the places the data could not resolve.

Core genome. The genes present in every member of a species — in practice, in ninety-nine per cent of them, because assemblies are imperfect and demanding a hundred per cent measures your worst assembly rather than the biology.

Coverage. How many times, on average, each position in a genome was read. Also called read depth. Low coverage does not announce itself; it produces confident wrong answers.

Dendrogram. A branching diagram of similarity. Distinguished from a phylogeny by not claiming to represent descent, a distinction routinely lost on the people reading it.

Dry lab. The computational side of the work. As against the wet lab, where the organisms are.

Efflux. A resistance mechanism in which the cell pumps the antibiotic back out. Invisible to a method that looks only for resistance genes it already knows.

FASTA. A text format for sequences: a header line starting with >, then the sequence. Designed in the 1980s, never formally specified, universal.

FASTQ. FASTA with a quality score for every base. The format a sequencer produces and the thing most pipelines start from.

GISAID. A database of viral genome sequences with an access agreement attached, established for influenza and central to SARS-CoV-2 data sharing. Its terms are the subject of a long argument about what open means.

GTDB. Genome Taxonomy Database. A taxonomy built from genome sequence rather than from the historical literature, which disagrees with the official one in places and is widely used anyway.

Homopolymer. A run of the same base — AAAAAA. The classic failure mode for several sequencing technologies, which lose count.

Isolate. A single organism grown up in pure culture from a sample. The unit most of this book is about.

k-mer. A substring of length k. Chop a sequence into every overlapping window of that length and you can compare, classify or assemble without ever aligning anything. The unit most assemblers and classifiers actually work in.

Kitome. The bacterial DNA that arrives in the extraction kit rather than in the sample. Reliably mistaken for a discovery when the sample itself contains almost nothing.

Lineage. A line of descent. In outbreak work, a group of isolates sharing a recent common ancestor closely enough to be treated as the same thing spreading.

Locus. A defined position in the genome. Plural loci.

MAG. Metagenome-assembled genome. A genome assembled out of a mixed sample without ever culturing the organism it belongs to.

Metagenomics. Sequencing everything in a sample at once, without culturing anything first, and working out afterwards what was in there.

MIC. Minimum inhibitory concentration. The lowest concentration of a drug that stops the organism growing — a number, as against the susceptible/intermediate/resistant call derived from it.

MLST. Multi-locus sequence typing. Strain typing from a handful of housekeeping genes, each reduced to an integer, the integers together making a profile and the profile getting a number of its own — the sequence type.

MLVA. Multi-locus variable-number tandem repeat analysis. A typing method based on counting repeated sequence units.

N50. The contig length at which half the assembly sits in contigs of that length or longer. The most-quoted assembly statistic and, on its own, close to meaningless.

Operon. A set of genes transcribed together as a unit.

Pangenome. The union of all genes found across a set of genomes, conventionally split into the core genome and the accessory genome. Where the line falls is a threshold somebody chose.

Panmixia. A population so thoroughly recombined that discrete lineages do not exist in it. The opposite pole from clonality, and the reason MLST needed the argument settled before it could mean anything.

PFGE. Pulsed-field gel electrophoresis. Cut the genome with a rare-cutting enzyme, run the fragments out on a gel with the field direction alternating, compare the resulting band patterns. The dominant outbreak-typing method for two decades, and a photograph rather than a number.

PHA4GE. Public Health Alliance for Genomic Epidemiology. An international consortium that produces open specifications for the contextual data attached to pathogen sequences.

Phage. A virus that infects bacteria. Frequently found integrated into bacterial genomes, where it is one of the things making two isolates of the same species disagree.

Phylogeny. A tree representing descent from common ancestors. Distinguished from a dendrogram by making that claim.

Plasmid. A DNA molecule separate from the chromosome, replicating independently, often carrying resistance genes and often moving between cells.

Posterior. In Bayesian inference, what you believe after seeing the data. As against the prior, which is what you believed before, and which you are obliged to state.

Preprint. A paper posted publicly before peer review.

Prior. See posterior. The thing critics of Bayesian methods object to and practitioners point out you were assuming anyway.

Provenance. The record of where a piece of data came from and what has been done to it.

QC. Quality control.

Recombination. The exchange of DNA between organisms rather than its inheritance from a parent. It breaks the assumption that a tree can represent the history, which is why so much of this book is about it.

Reference genome. One genome designated as the coordinate system everything else is described against. Choosing it is a decision with consequences that outlive the person who made it.

Serovar (also serotype). A variant distinguishable by the antibodies that bind to it. Historically how Salmonella was named, which is why Salmonella serovar names look like place names — because they are.

Sketch. A small summary of a genome — typically a sample of its k-mers — that supports approximate comparison thousands of times faster than the exact calculation.

SNP. Single-nucleotide polymorphism. A single-letter difference between two sequences. Pronounced “snip”, counted obsessively, and only meaningful relative to a stated reference and a stated set of filters.

Spoligotyping. A typing method for Mycobacterium tuberculosis based on which of forty-three specific spacer sequences are present. The result is a pattern, and some people can read the patterns from memory.

Surveillance. The routine sequencing of pathogens in the absence of a specific outbreak, so that an outbreak can be recognised when it happens. The part nobody funds in the quiet years.

Tagmentation. A library-preparation step that fragments DNA and attaches adapters in one reaction. Cheap, and sensitive to how much DNA you started with.

Transposon. A sequence that moves itself around the genome, frequently carrying other genes with it.

Type strain. The specific deposited culture that a species name is formally attached to. The requirement that one exist is what excludes most organisms on Earth from being named at all — the problem SeqCode was written to solve.

Wet lab. Where the organisms, reagents and pipettes are. As against the dry lab.

Workflow manager. Software that runs the steps of an analysis in the right order, on the right machines, and can resume where it stopped. Nextflow, Snakemake, WDL and CWL are the ones this book argues about.

About the Authors

Nabil-Fareed Alikhan is a microbial bioinformatician. He wrote the BLAST Ring Image Generator, which is the tool a great many people in this field have used at least once without ever learning who made it, and he has spent years on EnteroBase, which took Mark Achtman’s typing databases and turned them into something a laboratory anywhere can query. When scientific Twitter fell over in late 2022 he and a collaborator stood up a Mastodon server in a couple of days, expected fifty people, and got several thousand.

Lee Katz works in bioinformatics in United States public health, where the output of a pipeline is not a figure in a paper but a line in a report about somebody’s Salmonella. He wrote Lyve-SET and Mashtree. Much of his career has gone on the unglamorous half of the problem: making a method fast enough, and stable enough, that somebody can defend the result years later in front of people who are not impressed by any of it.

Andrew J. Page has worked on pathogen genomics at the Sanger Institute and the Quadram Institute, through a national pandemic response, and now at Origin Sciences on cancer diagnostics — which, as he put it on air, is “so not even microbes”, a sentence with fifteen years of microbial genomics behind it and no distress in it at all. He wrote Roary, SaffronTree, Socru and a long tail of other tools, and is a co-author of the argument in chapter two that most of them should probably never have been written.

The three of them met at conferences and at the microbial bioinformatics hackathons, tried to run a virtual lab talk series, could not get overworked scientists to volunteer to give the talks, and started a podcast instead. It has run for seven years and a hundred and fifty-six episodes.

Colophon

The audio was transcribed with Whisper large-v3 and the transcripts corrected against a vocabulary of the tool names, organisms and surnames that speech recognition reliably destroys. Names are spelled as the episode notes spell them, never as the transcript guesses them: left to itself the model renders one of the authors as a plausible-looking name belonging to nobody, a hundred and forty-four times.

Every quoted span in this book was checked against the corrected transcripts by a tool that fails the chapter if a quotation cannot be found. Punctuation inside quotation marks is the transcriber’s guess and not the speaker’s, so the check ignores it and compares words only.

Eight quotations in the book drop a word: a stutter or a false start, of the “whatever point you wanted to make, it’s, it is gone” kind. Nobody means those words, and printing them makes a fluent speaker read as though they were struggling. The tool allows the deletion and no more — it will not accept a quotation with a word added, changed or moved, and it requires the words that remain to sit close together in the recording, so that two remarks minutes apart can never be joined into one sentence somebody never said. Every such quotation is listed when the chapter is checked, and each was read back by a person.

The tool also refuses a set of phrases the podcast never used, which is a cheap way of keeping the register honest.

Corrections

Some of this is going to be wrong. A book assembled out of seven years of people talking from memory, checked against a machine transcript of that talking, will contain errors that nobody in the room noticed at the time and nobody has noticed since — a misremembered year, a method credited to the wrong group, a name spelled as the episode notes spelled it rather than as its owner does.

If you find one, say so at microbinfie.github.io. Corrections are listed there with the date they were made, and folded into the next revision of every format. Several corrections from listeners are already in this book, and at least one of them is in a chapter that would otherwise have been wrong.

Set in Georgia. Typeset from Markdown via pandoc. The cover is set in the same face, over a FASTA header — a format nobody designed, described in chapter 4.

Nobody wrote this down. So we did.

Index

References are to chapter numbers, not pages, so that they hold in every format this book is published in.

A

Antibiogram, 5

Antimicrobial resistance, 17

ARTIC network, 27

Assembly, 1, 6, 7, 8, 14, 15, 16, 18, 20

Average nucleotide identity (ANI), 11, 12

B

Bactopia, 20

Basecalling, 25

Bayesian inference, 10, 17

BLAST, 1, 5, 9, 11, 29

BRIG, 9

C

C and C++, 3

Campylobacter, 17

Careers, 2, 3, 5, 6, 9, 12, 15, 20, 26, 27, 28

CDC, 13, 29

CLIMB, 27

Cloud computing, 2, 19, 20, 23, 27

Conferences, 17, 22

Contamination, 1, 11, 12, 13

Core genome, 6, 14

Coverage, 1, 7, 12, 19, 28

D

Data sharing, 13, 22

Documentation, 27

E

EnteroBase, 6

Escherichia coli, 5, 6, 7, 11, 13, 15, 16

F

File formats, 3, 4, 9, 14, 18

Funding, 2, 19, 23, 26, 28

H

Haiti, 18

K

k-mers, 1, 7, 11, 12, 18

Klebsiella, 16

Kraken, 21

L

Large language models, 25

Listeria, 1, 5, 11, 12

M

Maintenance of software, 2, 20, 26

Mash and MinHash, 11, 25

Metagenomics, 11, 18, 19, 28

Mobile genetic elements, 5, 7, 14, 15, 16, 17, 18

Multi-locus sequence typing (MLST), 1, 2, 5, 6, 11

Mycobacterium tuberculosis, 9, 10, 15, 16, 17

N

N50, 7

NCBI, 2, 4, 7, 9, 12, 14, 18, 24

Neisseria, 6

Nomenclature, 8, 12, 13

O

Open source, 2, 11, 14, 16, 21, 26

Outbreak investigation, 1, 8, 11, 17, 18, 19, 22, 28

P

Pangenomes, 14

Peer review, 7, 10, 12, 24

Perl, 1, 3, 14, 20, 21, 25

Phylogenetics, 6, 9, 10, 11, 12, 13, 14, 15, 17, 19, 22, 29

Prokka, 2

Public health laboratories, 11, 18, 19, 20, 21, 23, 26, 28

Pulsed-field gel electrophoresis (PFGE), 5, 9

Python, 1, 3, 20, 21, 25, 28

R

R, 3, 9

Recombination, 6, 9, 29

Reference genomes, 5, 6, 17

Reproducibility, 4, 20, 21

Roary, 14

Rust, 3, 21

S

Salmonella, 1, 5, 6, 7, 11, 13, 14, 15, 16, 17, 22

samtools, 8

SARS-CoV-2, 18, 19, 22, 23

Sepia, 21

Sequencing technologies, 1, 4, 7, 8, 15, 18, 19

Single-nucleotide polymorphisms (SNPs), 8, 9, 16, 17, 18, 19, 22

Species definition, 12

Spoligotyping, 17

StaPH-B, 21

Staphylococcus aureus, 5

Surveillance, 15, 17, 18, 19, 22

T

Taxonomy, 12, 18, 21

The Gambia, 8, 17, 23

Training and teaching, 1, 3, 5, 21, 23, 27, 28

V

Velvet, 7

Vibrio cholerae, 12, 13, 18

W

Workflow managers, 20

Bold references mark the chapter that discusses the entry at length.