Dong Wang at the University of California San Diego and Steven Benner at the Foundation for Applied Molecular Evolution in Alachua, Florida, the two senior authors who between them designed the eight-letter alphabet and the transcription experiments, have shown that Escherichia coli RNA polymerase transcribes a DNA template carrying a synthetic letter at a rate the authors call comparable to a natural pair. Their paper appeared on 2 September 2026 in Nature Communications, pairing gel assays with four cryo-electron microscopy structures collected at the Stanford-SLAC Cryo-EM Center, at resolutions of 2.42 to 2.75 angstroms. It fills a gap left open since the alphabet was announced: the P-with-Z pair had been shown working with simple, single-subunit enzymes, and its behaviour in the many-part enzyme that cells themselves use for transcription had never been characterised.
The alphabet is called Hachimoji, from the Japanese for eight letters. Benner and colleagues introduced it in 2019, and it consists of the four natural bases plus four more, arranged into two extra pairs, P with Z and B with S. Each new pair rearranges the hydrogen-bond donors and acceptors on the base while keeping the overall shape of a Watson-Crick pair, which is what lets an enzyme built for four letters accept six or eight, provided the fidelity holds up as well as the geometry.
Each template held exactly one synthetic letter
The biochemistry runs on short DNA and RNA scaffolds, each carrying a single unnatural base at the position the enzyme is about to read. The team built four of them, one each for dP, dZ, dB and dS, offered all eight nucleotide triphosphates to each in turn, and quantified the resulting 32 template and substrate combinations from gel band intensities at the 15-second mark.
In every scaffold, adding the correct partner produced a clear extension band within 15 seconds. Single-turnover kinetics put the rate constants for dZ paired with incoming PTP, and for dP paired with incoming ZTP, at roughly two-fold below the natural dG and CTP pair measured under the same conditions, which the authors call comparable rather than equal. Chase experiments then showed the polymerase carrying on past the synthetic pair without stalling at the next position.
None of this happened inside a living organism, and any description of cells processing synthetic genetic information runs ahead of the experiment. Every measurement was made on elongation complexes reassembled from purified E. coli RNA polymerase and those scaffolds in buffer, the gel assays put together at 30 degrees Celsius and the structural samples at room temperature with far more nucleotide. The structural samples were then frozen one step before the chemical bond forms, by using an RNA that cannot be extended. Most of the gel work rests on two independent replicates.
The eight-letter claim is likewise assembled rather than run in a single tube. What is new here is the P:Z pair; the B:S pair had already been structurally characterised in the same enzyme by much of the same group in 2023, and each pair was tested on its own template. Orthogonality itself was prior art, demonstrated with the single-subunit T7 enzyme. What the two papers add is that it survives in the multi-subunit one: each synthetic letter strongly prefers its own partner, with a short list of slow minor mismatches the authors name, among them one synthetic letter pairing with another, dS with incoming ZTP.
Guanine slips in when Z loses a proton
Two combinations misbehaved, and one of them organised the rest of the paper. Within the P:Z system, a dZ template accepted GTP far more readily than any other non-cognate substrate. The other standout, a dB template accepting UTP, this group characterised earlier and traced to a rare tautomer of isoguanine.
Z carries a nitro group in the major groove. That group pulls electrons away from the ring and drops the base’s pKa to about 7.8, so at physiological pH deprotonation is favoured, and the deprotonated form’s hydrogen-bonding face mimics cytosine. Guanine pairs with it happily, in Watson-Crick-like geometry. The behaviour was already known from DNA replication and from PCR, where the fidelity of Z:P varies with pH, and the new results extend it to a polymerase whose active-site checkpoints, the authors note, are insufficient to overcome it. The checkpoints are not inert: correct incorporation still runs substantially faster than the mismatches, which is what makes the discrimination work at all.
The fix is a redesigned letter. The team used a recently developed analogue, 2′-F-α-carboxamide-Z, which they call Z*, replacing the nitro group with a carboxamide and adding a fluorine at the 2′ position to manage tautomerism. That raises the pKa above 10 and strongly disfavours deprotonation at neutral pH. With Z* on the template, GTP misincorporation is substantially reduced, which is the paper’s word for it, and not eliminated. Run the other way, incoming ZTP left extra misincorporation bands where Z*TP left none: on the dP template it paired correctly and then ran on, misincorporating against A and G at the positions after it, and on the dG template it mispaired at the templating position itself. On a further scaffold, dS, Z*TP still misincorporated, only less than ZTP did.
Those error rates need care, and the authors supply the caution themselves. Their single-nucleotide assays present one substrate at a time, with no correct nucleotide competing for the same site, which they describe as forcing the misincorporation; the frequencies observed that way are, in their words, likely overestimates of the actual error rates when the cognate substrate and all the others are present and compete simultaneously.
The nitro group cuts both ways
Across all four structures, in both the original and the redesigned form of the pair, the incoming letter sits in the active site in canonical Watson-Crick geometry with the template letter. Two of the four caught the trigger loop still open, and the pairing is already correct in both, which is how the authors know that geometry forms early and without help from the closed loop. In the two where the loop has closed it folds as it does on natural substrates, though less crisply in the Z* one, and in the sharpest of the four, the 2.42-angstrom dZ:PTP complex, the two catalytic magnesium ions take up the arrangement seen in native polymerases. That is the paper’s headline result, and it is the unremarkable one.
The structures also turned up something the fidelity story does not predict, and the kinetics are what carry it. Under otherwise identical conditions, a dZ template takes up incoming PTP measurably faster than a dZ* template does, which points the rate difference at the nitro group itself. The slowdown stops there. The authors report no measurable difference in the kinetics of the next nucleotide added, so the penalty sits on the incorporation step at the synthetic position and does not carry into the rest of the transcript.
The proposed cause is a single water molecule, placed automatically during model building, kept only if it cleared a map-quality score in the full map and in both half-maps, and then analysed geometrically. In the 2.42-angstrom dZ:PTP structure it sits 2.62 angstroms from the nitrogen of the dZ nitro group, offset by only 0.64 angstroms from the axis running perpendicular to the nitro plane, and it hydrogen-bonds to the backbone of two bridge-helix residues, A787 and A791. Nitro groups generate a positive electrostatic region above that plane, and the water’s distance and its offset from that axis are, the authors argue, characteristic of a lone pair donated into it and not of an ordinary hydrogen bond. Measured against the straight bridge helix of a published complex with no incoming nucleotide, the helix here is bent by about 2 angstroms at A787, the water sits at the hinge, and the authors read it as helping to hold that bend and so favour folding of the trigger loop, the mobile element that clamps down once a correct nucleotide arrives.
Sorting the particles that survived classification by trigger-loop shape shows what the paper calls the opposite trend between its two datasets: 76.8 percent folded shut in the dZ and PTP dataset, 19.4 percent in the dP and Z*TP one. Both figures are shares of a particle set, and by the paper’s own arithmetic each is a share of its own dataset rather than of a common total. Each is the fraction left after the images were curated, picked and sorted in several rounds, not of everything the microscope recorded. That comparison carries less weight than the raw numbers suggest, and for three reasons. The two datasets were not collected alike: the dZ and PTP set merges 4,128 new images with 1,146 movies from the group’s earlier structure, 527 of them shot at a 60-degree stage tilt, against 12,400 untilted images for the other. The template base changes between the datasets, so the molecule missing an electron-poor group at the relevant position is not Z* at all, it is the dP now sitting on the template, whose C7 has nothing for a water to reach. The incoming nucleotide also changes at the same time, from a purine to a pyrimidine, which the authors themselves offer as an alternative contributor, and they are explicit that the water interaction is unlikely to be the sole determinant.
Even the folded class in the Z* dataset is less far along than the Z one. The second catalytic magnesium is too weakly resolved to model, and the first sits about 3.4 angstroms from the incoming nucleotide’s α-phosphate, a distance the paper calls incompatible with coordination.
Aptamers before proteins
The nearer use is RNA that does a job without being translated. The 2019 alphabet already produced a fluorophore-binding version of the Spinach aptamer, and six-letter AEGIS DNA aptamers carrying P and Z have been engineered to target liver cancer cells, so an expanded RNA version of the same trick is what the authors say the eight-letter system can now be built toward. They call Z* an initial step toward a fully optimised pair, with further refinement of the Z base still required.
Two disclosures belong with that. Every unnatural phosphoramidite except dB, and all the unnatural triphosphates, were bought from Firebird Biomolecular Sciences in Alachua, Florida, with dB coming from Glen Research; Benner and two co-authors work at the Foundation for Applied Molecular Evolution in the same Florida town. The paper declares no competing interests.
Making protein from an eight-letter message has not been done, and the figure in which the authors lay out the pathway marks that branch with what its caption calls transparent arrows to say so. In cell-free protein translation, ribosomes have been persuaded to accept a non-standard amino acid using the B:S pair alone, not the full alphabet, and B:S has its own residual fidelity problem in that isoguanine tautomer, a rare form of the base that mismatches with U.
The authors are also testing a sideways route around the B-with-U mismatch, and it does not involve fixing the letter at all. Because B can pair with either S or U, an expanded codon table could assign XXS and XXU to the same amino acid, in the way the natural code spends several codons on leucine. Redundancy would absorb that error rather than chemistry having to eliminate it, and those experiments, the paper says, are already under way.