CDSKIT (/sidieskit/) is a Python program that processes DNA sequences, especially protein-coding sequences. Many functions of this program are designed to handle DNA sequences using codons (sets of three nucleotides) as the unit, and therefore, edits the coding sequences without causing a frameshift. Input and output format names follow Biopython SeqIO, but transformed records can only be written to formats whose required annotations remain meaningful. In particular, translations and length-changing operations should normally use FASTA rather than FASTQ.
Published stable versions of CDSKIT are available from
Bioconda. The master branch can be
newer than the latest Bioconda package. For users requiring a conda
installation, please refer to Miniforge
for a lightweight conda environment.
conda install bioconda::cdskit
cdskit -h
pip install git+https://github.com/kfuku52/cdskit
Pretrained cdskit localize TargetP-style .pt models run on CPU, but they
need the optional machine-learning dependencies (torch and scikit-learn).
CDSKIT is not published on PyPI. To install the GitHub development version
with the ml extra, include the repository URL explicitly:
pip install 'cdskit[ml] @ git+https://github.com/kfuku52/cdskit.git'
See Wiki for detailed descriptions.
-
accession2fasta: Retrieving fasta sequences from a list of GenBank accessions -
aggregate: Extracting the longest sequences combined with a sequence name regex -
backalign: Back-aligning CDS from unaligned CDS + aligned proteins -
backtrim: Back-translating a trimmed protein alignment -
codonstats: Printing codon-aware per-sequence and aggregate codon-usage statistics -
degeneracy: Extracting aligned 0/2/3/4-fold degenerate nucleotide positions -
filter: Filtering CDS by sequence-level quality rules -
gapjust: Adjusting consecutive Ns to the fixed length -
hammer: Removing less-occupied codon columns from a gappy alignment -
intersection: Dropping non-overlapping sequence labels between two sequences files or between a sequence file and a gff file -
label: Modifying sequence labels -
longestorf: Finding the longest ORF by six-frame translation (+/- strands, 3 frames each) -
localize: Predicting targeting peptide classes (noTP,SP,mTP,cTP,lTP) or compatible multi-label localization models from CDS or protein input -
localize-learn: Training alocalizemodel from tab-separated data, or auto-downloaded UniProt entries (explicitlabels oruniprot_cctext inference mode) -
mask: Masking ambiguous and/or stop codons -
maxalign: Removing sequences to maximize codon-based alignment area (MaxAlign) -
pad: Making nucleotide sequences in-frame by head and tail paddings -
parsegb: Converting the GenBank format -
plot: Plotting aligned CDS summaries, codon-state maps, or nucleotide alignment views with consensus codon/AA and AA frequency logos using matplotlib (--mode summary|map|msa; default output is PDF, override with--format) -
printseq: Print a subset of sequences with a regex -
rmseq: Removing a subset of sequences by using a sequence name regex and by detecting problematic sequence characters -
split: Splitting 1st, 2nd, and 3rd codon positions -
stats: Printing sequence statistics -
translate: Translating CDS nucleotide sequences to amino acids -
trimcodon: Trimming aligned CDS codon columns by clean-codon fraction -
validate: Validating aligned CDS quality and reporting issues
CDSKIT is designed for data flow through standard input and output. Streamlined processing may be combined with other sequence processing tools, such as SeqKit, with pipes (|).
# Example
seqkit seq input.fasta.gz | cdskit pad | cdskit mask | seqkit translate | cdskit aggregate -x ":.*" > output.fasta
Commands that have independent record- or search-level work expose
--threads INT. Small metadata-only commands intentionally run serially.
--threads 1: single-threaded (default)--threads 2or larger: multi-threaded--threads 0: auto-detect available CPU count
For resource safety, cdskit caps worker counts at 64 by default. Set
CDSKIT_MAX_THREADS to change that ceiling.
CPU-bound record commands select process workers from total sequence workload
rather than record count. The default crossover is 16,000,000 residues and can
be tuned with CDSKIT_PROCESS_PARALLEL_MIN_RESIDUES. Vectorized translation is
processed serially because worker startup costs more than it saves.
Long options use snake_case consistently, for example --seq_file,
--out_file, --codon_table, and --seq_name_regex. Older compact names
remain accepted for backward compatibility, but print a deprecation warning
that names the replacement. Boolean values consistently accept
yes/no, true/false, on/off, and 1/0.
cdskit TSV files are UTF-8, tab-delimited, rectangular, and written with LF
line endings. Input tables must have a unique, non-empty header and all
columns required by the selected operation. Shared biological columns use
accession, sequence, organism_group, localization, peroxisome, and
fold_id. Peroxisome values are written as yes or no. Multi-part TSV
reports use one header, a schema_version column, and a section column
instead of concatenating tables with different widths. Versioned reports
currently use schema version 2 on every row. See the 0.24 migration
guide for the old-to-new CLI and TSV mappings and
the changelog for release details.
Lightweight JSON models only need the base installation. Pretrained TargetP-
style .pt models, including experimental peroxisome-head candidates, run on
CPU but require the optional machine-learning dependencies described above.
Train with a local TSV:
cdskit localize-learn \
--training_tsv train.tsv \
--seq_col sequence \
--label_mode explicit \
--localization_col localization \
--perox_col peroxisome \
--model_out localize_model.json
Train by auto-downloading UniProt data:
cdskit localize-learn \
--uniprot_preset viridiplantae \
--uniprot_query "keyword:Transit peptide" \
--label_mode uniprot_cc \
--seq_col sequence \
--localization_col cc_subcellular_location \
--uniprot_fields accession,sequence,cc_subcellular_location \
--uniprot_exclude_fragments yes \
--uniprot_out_tsv uniprot_download.tsv \
--model_out localize_model.json
Train a BiLSTM+attention model (PyTorch required):
cdskit localize-learn \
--training_tsv uniprot_download.tsv \
--seq_col sequence \
--seq_type protein \
--label_mode uniprot_cc \
--localization_col cc_subcellular_location \
--model_arch bilstm_attention \
--dl_seq_len 120 \
--dl_embed_dim 32 \
--dl_hidden_dim 64 \
--dl_epochs 10 \
--cv_folds 3 \
--model_out localize_bilstm.pt
Notes:
--uniprot_presetcan be used alone or combined with--uniprot_query.- If
--label_mode explicitis used with UniProt source,--uniprot_fieldsmust include both--localization_coland--perox_col. - Reports include overall
class_train_accuracyplus per-class values (class_train_accuracy_noTP, etc.). With--cv_folds, per-class CV accuracies are also reported (cv_class_accuracy_noTP, etc.).
Predict from in-frame CDS:
cdskit localize \
--seq_file cds.fasta \
--model localize_model.json \
--report localize.tsv
Predict from protein sequences:
cdskit localize \
--seq_file proteins.faa \
--seq_type protein \
--model localize_model.json \
--report localize.tsv
Compatible multi-label models can also be trained from the public DeepLoc 2.1
benchmark data with python -m cdskit.deeploc_benchmark. TargetP 2.0 comparison
helpers are available through python -m cdskit.targetp_benchmark. See the
localize wiki page
for supported localization, membrane, and sorting-signal labels plus benchmark
commands.
Run the repeatable CPU hot-path suite before and after performance changes:
python -m cdskit.benchmark_hotpaths
python -m cdskit.benchmark_hotpaths --fullThe dependency-light development loop runs unit and integration tests without loading optional ML frameworks:
uv sync --locked --no-dev --group test
uv run --no-sync python -m pytest -q tests/unit tests/integration -m "not ml and not subprocess"The committed uv.lock is the reproducible development and CI resolution.
Package consumers can continue to install cdskit with pip; the lock is not an
upper-bound policy for the library's declared compatibility ranges.
See TESTING.md for the ML, coverage, quality, package, and focused rerun commands, plus the test layout and fixture conventions.
There is no published paper on CDSKIT itself, but we used and cited CDSKIT in several papers including Fukushima & Pollock (2023, Nat Ecol Evol 7: 155-170).
This program is BSD-licensed (3 clause). See LICENSE for details.
