ESO detects hypermutable sites in engineered DNA sequences and optimizes them away with DNAChisel, while preserving the amino-acid translation and optimizing for host codon usage. This implementation follows the approach introduced in Menuhin-Gruman et al. (2022, ACS Synthetic Biology) - see Citation below.
Genes built from repetitive or duplicated DNA elements (a side effect of standard codon optimization, which tends to reuse the same "best" codon repeatedly) are prone to mutate away during propagation in a host organism, through mechanisms like replication slippage and recombination-mediated deletion. ESO detects these hotspots and asks DNAChisel to route around them while it optimizes, using the empirical mutation-rate model from the EFM Calculator (Jack et al., 2015, ACS Synthetic Biology).
1. Requirements: Python 3.11 or newer. Check what you have installed:
python --versionIf that prints something below Python 3.11, or fails with "command not found"/"not
recognized", install a current version from python.org/downloads
first (the installer's default settings are fine) before continuing.
2. Install ESO. In a terminal, cd into this folder (the one this README is in), then:
pip install .This can take a minute or two the first time. If it ends with a line that doesn't contain
the word error, it worked - skip to step 3. If you do see an error, check
Troubleshooting below before asking for help.
3. Try it on a real sequence. Put a FASTA file (a plain text file starting with a >
line, then the DNA sequence) in a folder by itself, e.g. my_sequences/gene.fasta
containing:
>my_gene
ATGGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTTAA
Then run:
eso-optimize --input-folder my_sequences --output-path my_results4. Check the result. Look inside my_results/gene/ - final_sequence.txt has the
optimized sequence, recombination_sites.csv/slippage_sites.csv list what was detected
and fixed. If the very last line printed to the terminal was Success!, it worked.
That's the whole loop. Everything below this point is reference material for going further (different host organisms, custom scoring, locking regions from editing, and so on) - not required to get a first result.
If you'd rather be walked through setup interactively instead of reading this file, try the
ESO Onboarding Assistant,
a ChatGPT assistant configured specifically for onboarding onto this tool (see
docs/custom-gpt-setup.md for its configuration).
-
'eso-optimize' is not recognized/command not found: eso-optimize(common on Windows right after a fresh install): the command was installed, but its folder isn't on your terminal's PATH yet. Use this instead - it always works, no PATH needed:python -m eso.cli --input-folder my_sequences --output-path my_results
(Every
eso-optimize ...example in this README works identically aspython -m eso.cli ....) -
An error mentioning
Microsoft Visual Studio,CMake, or "failed building wheel" duringpip install .: this meanspiptried to compile a dependency from source instead of using a prebuilt version - almost always fixed by upgrading pip first (python -m pip install --upgrade pip) and trying the install again, since an olderpipcan miss prebuilt wheels that a newer one finds. -
ModuleNotFoundErrororImportErrorright after install: double check you're runningpython/eso-optimizefrom the same Python installation you ranpip install .with. If you're not sure, runpython -m pip show eso- if that fails, you installed into a different Python than the one you're now running. -
"No such file or directory" for your input folder:
--input-folderis relative to wherever your terminal's current directory is - eithercdthere first, or use a full path (e.g.C:\Users\you\my_sequencesor/Users/you/my_sequences). -
Nothing in the output folder / an empty
resultslist: ESO looks for files ending in.fasta/.fna/.ffn/.faa/.frn/.fa/.gb/.gbk/.genbank(optionally gzipped), directly inside--input-folderor one level under it - check your file's extension matches one of those. -
Anything else: the error messages this tool prints are meant to be read directly and acted on (not just a Python traceback to decode) - if one isn't clear, that's a bug in the tool itself, worth reporting rather than working around.
- Replication slippage - short tandem repeats (a base unit of length 1-15 repeated 3+ times, or a single nucleotide repeated 4+ times) that polymerase can skip or duplicate during replication.
- Recombination-mediated deletion (RMD) - pairs of near-identical sites (16+ nt, within Levenshtein distance 1 of each other) that homologous recombination can delete between.
- Methylation motifs - sequence motifs recognized by host methylation machinery,
which can trigger repair-associated mutation. Given as a PSSM/MEME file, a bundled
common motif (E. coli's Dam/Dcm systems), or your own IUPAC consensus string
(e.g.
"GATC") - no file needed for either of the latter two.
Two independently-developed implementations of the recombination/slippage detectors are
included (eso.detection.recombination/eso.detection.slippage, and
eso.detection.staubility_variant) - they haven't yet been reconciled into a single
canonical algorithm. Both are routed through eso.detection.dispatch
(find_recombination_sites(seq, num_sites, mode="thorough" | "fast") and
find_slippage_sites(seq, num_sites, mode="default" | "fast")), also exposed via
eso.pipeline.main(..., recombination_mode=..., slippage_mode=...) /
eso-optimize --recombination-mode --slippage-mode. See
docs/detector-comparisons.md for the tradeoffs,
benchmarks, and bugs found while comparing them - recombination's two modes trade off
sensitivity for speed, but slippage's are fully equivalent in what they detect (a pure
speed choice). Methylation motif detection only has one implementation
(eso.detection.methylation) - a second was built and compared, but deleted after it
turned out to disagree in accuracy, not just speed (see the same doc for the writeup).
Not included: STABLES also carried a junction-linker hotspot checker
(linker_suspect_utils.py), applying the same slippage/recombination-detection idea to a
different, narrower problem - checking whether joining a target gene to a host sequence
via a linker creates a new hotspot at the junction. It was ported into an earlier version
of this repo (eso.detection.junction_linker) but removed: its detection thresholds (a
4-repeat homopolymer check, a 12-mer recombination check) don't match the calibrated -9
filter this codebase's other detectors use, and tracing them back into the original
STABLES select_fusion_linkers.py pipeline showed they act as a near-absolute veto in
STABLES' own linker-selection algorithm - deliberate, STABLES-specific conservatism for
screening many cheap, redesignable linker candidates, not a general-purpose hotspot
detector. That's a different problem from ESO's (evaluating one already-chosen sequence),
so it was judged out of scope for this library rather than reconciled with the other
detectors' thresholds.
poetry install
# or: pip install -e .Word-document diff reports (eso.report) need the optional docx-report extra:
poetry install -E docx-reportfrom eso import main
message, results = main(
input_folder="path/to/fasta_files",
output_path="path/to/output",
optimize=True,
mini_gc=0.3,
maxi_gc=0.7,
method="use_best_codon",
organism_name="e_coli", # or a TaxID, or a bundled table (see eso.codon_usage)
)Or from the command line:
eso-optimize --input-folder path/to/fasta_files --output-path path/to/output --organism-name e_coliFor each FASTA/GenBank file found in input_folder, this writes to output_path/<file_stem>/:
final_sequence.txt- the optimized sequence, plus CAI-before/after and edit-count stats.recombination_sites.csv/slippage_sites.csv/motif_sites.csv- detected hotspots.sequence_comparison.docx- a diff view of original vs. optimized sequence (if thedocx-reportextra is installed).
organism_name accepts anything supported by
python-codon-tables
(a species name or NCBI TaxID), or one of the bundled custom tables in
eso.codon_usage.CODON_USAGE_TABLES (C1, kompas, human_antibody_heavy_chain,
human_antibody_light_chain) for hosts not in that database.
See examples/antibody_optimization for a complete
worked example (human antibody heavy/light chain optimization).
main() (above) is file-in, file-out - convenient for a one-off CLI run, but if you
already have sequences in memory as part of your own code (e.g. generated, fetched from a
database, or produced by an earlier step in your own pipeline), you don't need to write
them to disk first. eso.optimize.optimization_engine (also importable as
eso.optimization_engine) takes a plain DNA string and returns a plain DNA string - no
files involved:
from eso import optimization_engine, suspect_site_extractor
seq = "ATG" + "GCT" * 15 + "TAA" # any DNA string you already have
# 1. detect hotspots (skip this step and the dataframes below entirely if you only
# want codon/GC optimization, with no hotspot avoidance)
sites = suspect_site_extractor(seq, compute_motifs=False, num_sites=50)
# 2. optimize, avoiding what was detected
final_seq, objectives_summary, num_edits = optimization_engine(
seq,
organism_name="e_coli",
df_recombination=sites["df_recombination"],
df_slippage=sites["df_slippage"],
)final_seq is a plain str you can feed straight back into whatever your own code does
next (write it out yourself, pass it to another function, etc.) - nothing here touches
the filesystem. suspect_site_extractor (also importable as
eso.suspect_site_extractor) is the same detection step main() runs internally; call
it on its own if you only want the hotspot dataframes, with no optimization at all.
This composes directly with a custom scoring function - just pass custom_score_fn=...
to optimization_engine instead of organism_name:
final_seq, _, num_edits = optimization_engine(
seq,
custom_score_fn=my_model.predict, # any function: whole seq (str) -> a number, higher = better
df_recombination=sites["df_recombination"],
df_slippage=sites["df_slippage"],
)See optimization_engine's docstring (eso/optimize.py) for every parameter
(mini_gc/maxi_gc, orf_regions/exclusion_regions, method, and so on) - everything
available via main()/the CLI is available here too, just without the file layer.
By default, optimization scores codon choices against a codon-usage table (CAI/tAI-style,
via organism_name). To score sequences with your own logic instead, write a Python file
defining a score(seq) function (seq is a plain DNA string like "ATGCGT...", the whole
ORF being optimized; return a number, higher = better) - copy
examples/custom_score_template.py as a starting
point.
From the command line:
eso-optimize --input-folder path/to/fasta_files --custom-score-file my_score.py(--custom-score-file overrides --organism-name/--method.) A --custom-score-file
with a mistake in it - a missing score function, a typo, a function that crashes or
returns the wrong type - fails immediately with a plain-English message, before any
optimization runs, rather than surfacing later as a Python traceback.
From Python, either call optimization_engine/eso.pipeline.main directly with
custom_score_fn/custom_score_minimize, or reuse the same file-loading + validation the
CLI uses via eso.custom_score.load_custom_score_from_file:
from eso.custom_score import load_custom_score_from_file
from eso.optimize import optimization_engine
score_fn = load_custom_score_from_file("my_score.py") # same validation as the CLI
final_seq, _, _ = optimization_engine(seq, custom_score_fn=score_fn)
# or, skipping the file entirely:
final_seq, _, _ = optimization_engine(
seq,
custom_score_fn=lambda whole_orf: whole_orf.count("G") + whole_orf.count("C"),
)custom_score_fn is called once on the whole ORF being optimized, on every trial mutation
tried during optimize() - this can be slow for a long sequence or an expensive function
(a warning is raised). custom_score_minimize=True (--custom-score-minimize on the CLI)
treats a lower custom_score_fn value as better, instead of higher.
Scope: custom scoring is automatically restricted to orf_regions (one scored region
per ORF, matching how the built-in CAI/tAI codon-usage scoring is already scoped via
DNAChisel's CodonOptimize(location=orf, ...)) - custom_score_fn never sees any non-ORF
flanking sequence (UTRs, locked/excluded regions). If you don't pass orf_regions, this
is the whole sequence (trimmed to a multiple of 3), same as everywhere else in ESO.
An earlier version of this feature also supported a "windowed" mode (scoring
fixed-size chunks and summing them, mirroring how the built-in CAI/tAI scoring works
internally) as a claimed speed optimization. It was removed after benchmarking found no
case where it was actually faster than the whole-ORF evaluation above - comparable at
best, meaningfully slower at worst, since DNAChisel's own optimizer ends up calling the
score function considerably more often when it's chunk-localizable - while carrying a
real, unpreventable correctness risk (a score that doesn't genuinely decompose per-chunk,
true of most real external/ML models, would silently compute a different, structurally
unrelated quantity, with no reliable way to detect this automatically). See
docs/detector-comparisons.md for the full investigation, including an initial benchmark
that was itself flawed and had to be corrected before the removal decision was made.
By default, the entire sequence is treated as one in-frame, translation-preserving ORF
with nothing locked. To instead give each sequence its own ORF region(s) (e.g. skip a
UTR) and/or exclusion regions that must never be edited (e.g. a known regulatory
element), pass indexes (from Python) or --indexes-file (from the CLI).
From the command line, point at a JSON file - copy
examples/indexes_template.json as a starting point:
[
{
"file": "my_gene",
"seq_index": "0",
"orf_regions": "1-6, 51-68",
"exclusion_regions": "1-6, 50-68"
}
]eso-optimize --input-folder path/to/fasta_files --indexes-file indexes.jsonfileis the FASTA/GenBank file's stem (no extension, e.g."my_gene"formy_gene.fasta).seq_indexis which record within that file, 0-indexed in file order, as a string.orf_regions/exclusion_regionsare 1-indexed, inclusive region strings (e.g."1-6, 51-68"for two separate regions); omitexclusion_regions(or use"") for no exclusions.
A malformed --indexes-file (bad JSON, missing file/seq_index) fails immediately with
a plain-English message; malformed region strings themselves are validated the same way
whether indexes came from a file or was passed directly to eso.pipeline.main/main.
From Python, pass the equivalent dict directly:
from eso import main
message, results = main(
input_folder="path/to/fasta_files",
indexes={("my_gene", "0"): ("1-6, 51-68", "1-6, 50-68")},
)Motif detection (--compute-motifs) isn't only about methylation - the same PSSM-based
scanner (eso.detection.methylation.find_motif_sites) works for any short sequence motif
you want to flag. Needs at least one motif, from any combination of three sources:
-
A MEME-minimal-format PSSM file (
--motifs-path/motifs_path=) - the original option, for a curated or experimentally-derived motif set you already have as a file. -
Bundled common motifs (
--common-motifs dam,dcm/common_motifs=["dam", "dcm"])- no file needed. See
eso/detection/common_motifs.py(COMMON_MOTIFS) for the full list and sources; currently: - Methylation (E. coli, already a first-class host here - see
eso.codon_usage's bundlede_colitable):dam(GATC, N6-methyladenine),dcm(CCWGG, C5-methylcytosine on the internal C). - Cryptic ribosome binding:
shine_dalgarno(AGGAGG) - flags a copy of the bacterial RBS consensus occurring inside a coding region, a known source of unintended internal translation initiation. - Cryptic bacterial promoter elements:
sigma70_minus35(TTGACA),sigma70_minus10(TATAAT) - the two sigma70 hexamers; an accidental occurrence of either inside a coding sequence is a classic source of unwanted transcription. Each hexamer is flagged independently (a real promoter needs both, correctly spaced, which this detector doesn't check for) - treat an isolated hit as a coarse screen, not a confirmed cryptic promoter. - Not included, and why (see the module docstring for the full explanation):
transcription terminators (a secondary-structure property, not a fixed linear
motif), the Kozak sequence (something to match, not avoid - a different problem),
and restriction enzyme sites (already covered comprehensively by DNAChisel's own
dnachisel.list_common_enzymes()/EnzymeSitePattern, usable directly withAvoidPatternduring optimization - no need to duplicate it here).
- no file needed. See
-
Your own IUPAC consensus string - the easy custom-motif path, no PSSM/MEME file needed:
from eso.detection.motif_utils import motif_from_consensus, motifs_from_consensus_dict my_motif = motif_from_consensus("my_site", "GANTC") # N = any base, standard IUPAC codes # or define several at once: my_motifs = motifs_from_consensus_dict({"site_a": "GATC", "site_b": "CCWGG"})
Pass the result as
relevant_motifstoeso.detection.methylation.find_motif_sitesdirectly, or combine with the other two sources yourself before calling it - there's no CLI flag for this one yet (a MEME file is still the CLI's path for anything beyond the bundled common motifs).
Most published methylation/restriction motifs (REBASE, NEB's technical notes, the primary
literature) are naturally given exactly this way - a short consensus sequence with IUPAC
ambiguity codes (W = A or T, N = any base, etc.) - not as a position-probability
matrix, so this avoids hand-authoring a MEME file just to check for one. For organisms
beyond E. coli, REBASE is the standard reference database for
methylation motifs - look up the motif there and pass it straight to
motif_from_consensus.
poetry install --with dev
pytestIf you use this tool, please cite the paper it implements:
Menuhin-Gruman, I., Arbel, M., Amitay, N., Sionov, K., Naki, D., Katzir, I., Edgar, O., Bergman, S., & Tuller, T. (2022). Evolutionary Stability Optimizer (ESO): A Novel Approach to Identify and Avoid Mutational Hotspots in DNA Sequences While Maintaining High Expression Levels. ACS Synthetic Biology, 11(3), 1142-1151. https://doi.org/10.1021/acssynbio.1c00426
@article{menuhingruman2022eso,
title = {Evolutionary Stability Optimizer (ESO): A Novel Approach to Identify and
Avoid Mutational Hotspots in DNA Sequences While Maintaining High
Expression Levels},
author = {Menuhin-Gruman, Itamar and Arbel, Matan and Amitay, Niv and Sionov, Karin
and Naki, Doron and Katzir, Itai and Edgar, Omer and Bergman, Shaked
and Tuller, Tamir},
journal = {ACS Synthetic Biology},
volume = {11},
number = {3},
pages = {1142--1151},
year = {2022},
doi = {10.1021/acssynbio.1c00426}
}