2,002 Escherichia coli genomes · PanGBank 11587 · Bacformer

PanGramGraph

github.com/ggautreau/PanGramGraph

Every E. coli genome in PanGBank’s pangenome of the species, walked gene by gene from its replication initiator, and a genomic language model, Bacformer, asked at every position what comes next. Read in its own vocabulary it names the next gene nine times in ten, as often as the pangenome graph does, more often than the graph where the genomes part, and it hesitates where the pangenome is plastic.

The graph

hover a box · click to pin · drag to move · Shift + wheelwheel to scroll along · Ctrl + wheel to zoom
20
The 540 complete chromosomes, from dnaA, in genes. Drag or swipe the frame, or click anywhere, to open the window there. share of genes in a region of plasticity
expected by the model, follows it in no genome line width ∝ log genomes sharing it bars on top: the model’s perplexity at each column

Help

The graph

Left to right
position along the chromosome, counted in genes, not base pairs: from dnaA (column 0) in the dnaA window, from the window’s first gene elsewhere.
A box
one gene family of the PanGBank pangenome, at the column where it most often sits. Boxes stacked in a column: the genomes disagree there, the most common on top. That is a fork.
Colour
PPanGGOLiN’s partition of the species: orange persistent, green shell, blue cloud.
Solid line
an adjacency some genomes carry; thicker means more genomes (log scale).
Dashed box
a family the model expects right after the best-carried family of a column, and that follows it in no genome. Two per column; Show what the model expected hides them.
Genomes carrying
the slider hides families carried by fewer genomes.
Bars on top
the model’s perplexity at each column, on the family actually there in each genome: 2 to the power of the mean surprise, −log2 P. 1 means always certain and right; 2 is as unsure as a coin toss between two families. Hover a bar for its detail, including 2entropy, the number of families the model hesitates between. P is read through the decoder (see Exploring); a family none of the model’s 15 best clusters reads as (1.5% of calls) gets the probability of the 15th, so the value errs on the low side.
Fewer genomes on the right
draft assemblies stop where their contig stops: in the dnaA window all 2,002 genomes reach column 3, 1,359 reach column 79.

Exploring

Hover a box
a card opens beside it: the family across the genomes, what follows it, and the model’s call after it, averaged over every genome that carries it.
Click a box
pins the card and computes the model’s exact call in one genome, with a search for any family at that position. Draw from this path starts a path from that genome’s own proteins. Needs the local server.
Pick a genome
the genome box lists, before you type, genomes that carry the family and well-known strains. Type any mix of strain, serotype (O157, H7), sequence type, host, isolation source, country, year or accession: every word must match (human urine, shigella, cattle 2012). ↑ ↓ move, Enter picks, Esc gives the previous choice back. Beyond the dnaA window drafts are greyed out: only the complete chromosomes are read there.
Probabilities
the softmax of the model’s output over its 50,000 clusters, as percentages. They are not p-values.
Decoder
the model predicts its own clusters, not PanGBank families. A cluster is read as the families actually present when the model predicts it, P(family | cluster), learned on one half of the genomes and applied to the other. An exemplar bank of 9,814 clusters, used before, misses many of them. Clusters read as the same family share one row, their probabilities summed (4 clusters).

Drawing a path

Draw
switch the graph to Draw a path and click boxes in order, or press on a box and slide across the next ones in one stroke, as on a phone keyboard; the trace follows the pointer. Slide back to erase; slide up or down a column to change the choice there.
The model’s call
for the family after exactly that sequence of proteins, when you release. Every box glows by how likely its family is to come next; the 12 likeliest carry their probability. Any order is allowed, including one no genome carries: that is how to ask the model about a combination it has never seen.
Edit
drag a box of the path onto another box to change that step, or off the graph to remove it; Esc cancels. Undo reverts the last change, a whole stroke included.
A genome’s path
pick a genome in the Genome box, then click a box: the path becomes that genome’s own proteins up to it (Take a genome’s path up to the box). In a window beyond dnaA the model also reads, unseen, the 400 proteins before the window in that genome.
Attention
the Attention button opens the Attention tab on the path.

Along the chromosome

The map
under the bar, the complete chromosomes from dnaA, in genes, shaded by the share of genes in a region of plasticity. The frame is the window shown: drag it along, swipe it sideways with a finger or two on a trackpad, or click anywhere; the window opens where the frame is left. With the map selected, ← → move one window, PgUp PgDn 500 genes, Home goes back to dnaA.
The bar above the graph
Previous and Next (keys P and N) move 60 genes at a time: a window shares 20 columns with the next. The search box suggests every gene name of the complete chromosomes, with its place; it also takes a PanGBank family or a position from dnaA (1500). A gene already in the window is shown there, otherwise its window opens. dnaA window returns here. A link to a place: pangramgraph.html#at=lacZ.
Which genomes
beyond the dnaA window only the 540 complete genomes are in one piece and have the model’s call at every gene, so a window elsewhere is read in them. It starts at a backbone family, present once in at least 95% of the chromosomes, and walks 80 genes in dnaA’s direction. The local server builds it in about two seconds. The text and tables of Findings describe the dnaA window.

Attention

What is read
the path drawn in the Graph tab, and, for a genome’s own path in a window beyond dnaA, the 400 proteins before it (one column, summed). Without a path, the tab offers the first ten genes of the window.
The heads
Bacformer has 12 layers of 8 heads. The map gives one cell per head, coloured by how far back it looks on average among the path’s proteins, or by its share on the previous protein, the start token or the protein itself. Click a cell to show that head’s matrix.
The matrix
rows are the path’s proteins, columns what each one attends to; causal, so nothing above the diagonal. Leave the start token out hides the attention sink and renormalises each row over the rest. Colour on a square-root, linear or log scale.
The last row
for every layer and head: where the model looks when it calls the next family.
How far to read it
attention shows where the model looks, not why it calls what it calls; heads that rest on the start token are often idle, not blind. Computed live by the local server.

Coupled regions

What is shown
accessory gene families in different insertion spots whose presence goes together (arcs above the chromosome) or apart (dashed, below) within lineages and within tertiles of accessory content, beyond a permutation null. An element is a cassette of an insertion spot that comes and goes as one; a set is a community of co-occurrence links. These tests read the presence of genes in the 540 complete genomes, not the model.
Two layers
Long range: elements more than 500 genes apart (533 links). Close range: 50 to 500 genes apart in the genomes carrying both, in separate insertions (113 links), where the model’s use of the context is strongest; a pair on both sides of dnaA is drawn the short way round, leaving the map at both ends. Arcs join the elements’ median positions; a link’s detail also gives how far apart the two sit in the genomes carrying both (when fewer than 10 genomes do, as for most close-range avoidance links, only the median positions). The map, the three lists (sets, links, elements) and the detail work on either; the lists are also the way in on a phone and from the keyboard (arrow keys, Enter).
Colours and lines
with model verdict, an arc’s colour is what the model says about the link (below), and the elements are grey; with set, the set of its two ends. The line says what the link is in the genomes: solid co-occurrence, dashed avoidance, dotted one element at several spots. Click a legend entry to hide or show its arcs; the legend counts the links the filters keep.
What the model says
an in-silico knockout of Bacformer on the complete chromosomes, in full precision: in each genome carrying both elements (or one, for avoidance), the element that comes first from dnaA is deleted and the model’s probability of the other element’s first family is read again, against deletions of other whole accessory elements of about the same size and distance, with each genome read by a decoder learned on the other lineages. knows: the other element moves the way the genomes go, more than under control deletions, in both halves of the lineages, and more than the other elements the same deletion moves (specificity, in z and in raw units). not specific: it moves the way the genomes go, but no more than everything else downstream; the deletion shifts the model’s expectation broadly. opposite: a specific move against the genomes. none: no move that passes the tests either way; the detail says which condition failed. not testable: fewer than 5 lineages, no readout, only partial deletions, or changes below the numerical floor.
How far to trust each verdict
knows is the strongest claim: calibrated on pseudo-deletions and corrected over the links, and it holds in two lineage-disjoint halves. not specific shows the model uses distant content, not that it knows the pair. none is only as strong as its smallest detectable effect and its distance: the model’s use of context decays (a 2-gene deletion moves the call of a gene 10 genes on by 0.044 in logit, 100 genes on by 0.014, 1,000 genes on by 0.006), so a none tested beyond 500 genes means the test could not see it, not that the model ignores the link. not testable says nothing. A pilot on 24 long-range links, tested at a median 1,414 genes, found none known, 3 not specific, 9 none and 12 not testable. The detail of a link gives every number: genomes tested, lineages, the median effect overall and in each lineage half, the share tested within the model’s range, the smallest detectable effect and the q values each verdict was decided on.
A correction
an earlier version of this tab called a link replicated when it held in two halves of the genomes, but those halves shared lineages (6 of the 8 lineage clusters were in both). With lineage-disjoint halves, 267 of the 533 long-range links replicate; only links replicated in lineage-disjoint halves shows those (on by default), and the sets list greys the sets left with none. Every close-range link was selected that way.
One element at several spots
when, in the genomes carrying both, a copy of each sits in one spot, the “link” is one element that lands at different places in different genomes, not two regions that go together (110 long-range links, 4 close-range ones). Flags in a link’s detail also say when the two lie under 100 genes apart (possibly one transfer tract), in separate insertions, or share an insertion site.
Without an element
click an element, pick a genome carrying it at its spot, and Delete it and read the model: the page’s server deletes it, then 8 whole accessory elements of about its size in turn, and draws the mean change of the model’s expectation per 25 genes out to 1,500 genes downstream, against the controls’ range. A close-range element is deleted as an element of the close-range sets, and its links are the close-range ones. specific: some elements downstream move differently from the others at their distance (q < 0.05); broad: none does, while the deletion moves the model more than a typical one; like a control deletion: none does, and it moves the model no more than a random element of its size. Each linked element is then placed: read, with its change and whether it goes the way the genomes go, or before the element (the model reads from dnaA on), beyond 1,500 genes, not at its spot, or absent. A few seconds on a GPU; on a CPU, as on the public Space, 2 to 4 minutes, with the progress shown; a computation keeps running if you look at something else meanwhile.

Mouse and keys

Move
drag the background, or drag anywhere with the wheel pressed. Shift + wheel scrolls along the graph; in full screen the wheel alone does.
Zoom
Ctrl + wheel, or the buttons: Overview fits the whole window, Fill height fills the height and scrolls sideways, Full screen gives the graph the screen.
Keys
+ − zoom, 0 back to 100%, F overview, H fill height, N P next and previous window, Esc closes a card, cancels a drag or leaves full screen, ? opens this help.

The local server

Start it
python3 serve_live.py in origin-walk/, then open http://localhost:8765/. It listens on 127.0.0.1 only.
What needs it
the model’s call in one genome, drawn paths, windows beyond dnaA, attention maps and what the model does without a coupled element. The graph of the dnaA window, the coupled regions with their verdicts and the findings work without it. A page opened from another local server (another port) reaches it too.