Method
How the scan works.
Two detectors, the same code in every browser, and a fingerprint of every result.
What this reruns
In September 2026, Yoon et al. at Anthropic reported a campaign of 949 AI agent sessions that searched about 1.9 billion protein clusters for genes working alongside reverse transcriptases. One worker read the raw DNA upstream of a retron-like RT in a jumbo phage and noticed the same short sequence repeated again and again. That led to ART, a new family of array-associated reverse transcriptases, with repeat arrays in 28 of 95 family members.
Nature's coverage calls the hosts giant viruses. The preprint says jumbo phages: bacteriophages with genomes over 200 kb. Every named genome is a Staphylococcus or Listeria phage. So this catalog scans every complete jumbo phage in GenBank first, and the giant viruses second.
The agents in the preprint were language models. The agents here are browsers running one deterministic program, so any two of them must reach the same bytes, and a disagreement means something.
Two detectors
CRISPR mode
Short repeats with CRISPR-sized spacers, following CRT (Bland et al. 2007) with MinCED's defaults. An exact 8-base seed shared by the copies, repeat edges grown while 75% of copies agree, every copy of the consensus collected within 20% mismatch, then filters: repeats of 23 to 47 nt, spacers of 26 to 50 nt within 12 nt of each other, and spacers no more than 62% similar to each other or to the repeat.
ART mode
No CRISPR finder admits ART's spacers, which run 120 to 220 nt. This detector follows the preprint's own recipe: an exact 12-base word, with at least three kinds of base and no run of six, that recurs at least three times at a regular spacing of 100 to 450 nt (30% tolerance). Seeds that run in tandem closer than 100 nt are excluded. Edges grow while 75% of copies agree, up to 60 nt. Copies with a mutated seed are rescued at 80% identity beyond the ends and 70% inside a gap, and adjacent spacers must be no more than 62% alike.
Labels
A repeat finder sees many things that are not arrays. Every call gets one label, from where it sits:
- ART-like: a long-spacer array outside any gene, upstream of a reverse transcriptase candidate, with no gene of 300 nt or more between. That is the adjacency rule the preprint used.
- CRISPR-like: a CRISPR-sized array outside any gene.
- Intergenic array: outside genes, but with no RT beside it.
- Coding repeat: most copies sit inside a protein-coding gene. Usually a repeated protein motif, such as a collagen-like stretch.
- RNA gene cluster: overlaps tRNA or rRNA genes, which repeat in tandem.
A label says where an array is. It says nothing about what the array does.
Finding the reverse transcriptase
GenBank names most RTs ("RNA-directed DNA polymerase", "retron-type", "group II intron"). But in three of the seven Staphylococcus ART phages the RT is only a "hypothetical protein". So a gene with no name, 350 to 900 amino acids long and carrying the catalytic [YF]xDD motif, also counts. That motif is common, around a dozen such genes per jumbo phage, so an unnamed gene counts only within 500 nt of the array. The seven known loci sit 274 to 288 nt away. A named RT counts within 3 kb.
The real test is a protein profile search (Pfam RVT_1 and the myRT profiles, as in the preprint). That does not run in a browser yet, so the motif rule is a stand-in, and it is stated as one.
Controls
E. coli K-12 MG1655 (NC_000913.3) carries two textbook CRISPR arrays. The scan finds both: 13 copies of a 29 nt repeat at 2,877,701 and 7 copies at 2,904,014. It also finds a tRNA cluster, and labels it one.
The seven Staphylococcus jumbo phages the preprint names all come back ART-like. In MarsHill (MW248466.1), five copies of ATATGAATACGTAT start 213, 197, 210 and 173 nt apart, the spacings printed in the preprint's Figure 2, and the array ends 287 nt upstream of the RT the family was seeded from (QQM14740.1).
The preprint also names Listeria phage LPJP1 and Staphylococcus phage PB50 as family members. The scan finds no array in either. Only 28 of the 95 members had a detectable array, and the per-genome calls are drawn in a figure, so whether these two should have one is not settled here.
Replication
Every genome carries two SHA-256 fingerprints. One covers the sequence as uppercase bases, so everyone knows exactly which DNA was read. The other covers the result: the settings, the sequence fingerprint and every call with its label. The engine uses integer arithmetic only, so the result is identical on any machine.
When your browser scans a genome from the catalog, it fetches the sequence from NCBI itself, runs the same code, and compares. A match is an independent replication of the published result. A divergence means the two runs differ: a different sequence, a different engine, or a bug worth reporting.
A replication shows the published result is what this code gives on this sequence. It does not show that the arrays are real biology. That takes the bench.
Limits
- Arrays with fewer than three copies are missed, and so are copies more than 20% (CRISPR mode) or 30% (ART mode) away from the consensus.
- The metagenomic contigs where ART was first seen are not public, so they cannot be scanned.
- Replications are counted in each browser for now. A public, shared record of them is written but not deployed.
- Only complete genomes are in the catalog. Any other NCBI nucleotide accession can be scanned from the scanner, live.
Sources
- Yoon, Athukoralage, Ameisen, Kauderer-Abrams, Perry and Durrant. Autonomous AI agents discover reverse transcriptases with tandem repeat arrays, 2026. Also on alphaXiv.
- Nature news, 25 September 2026.
- Bland et al. CRISPR Recognition Tool (CRT), BMC Bioinformatics 2007, and MinCED.
- Sequences and feature tables from NCBI Nucleotide, fetched by your browser through E-utilities.