Huntingtin (HTT) is a large, evolutionarily conserved protein essential for embryonic development, intracellular transport, and neuronal survival. Near its N-terminus is a polyglutamine (polyQ) tract encoded by a repeating CAG codon; in humans, expansion of this tract beyond roughly 36 repeats causes Huntington's disease, a progressive neurodegenerative disorder. Most of the rest of the protein forms a huge HEAT-repeat solenoid that scaffolds interactions with dozens of other proteins.
What is in this dataset?
Click to flip
This dashboard combines HTT ortholog sequences pulled from UniProt and NCBI across hundreds of species, spanning chordates, nematodes, and arthropods. Each entry is annotated with taxonomic rank, full sequence length, and the length of its polyQ and polyP (polyproline) tracts, detected with an identical rule-based method across both source databases – see “Source Pipeline & Methodology” below.
What can I discover here?
Click to flip
Compare polyQ/polyP tract length across species and taxonomic groups, see which lineages retain or lack these repeat regions, view or download full protein sequences, and copy selected species as FASTA for alignment tools like Clustal Omega. Sort and filter the table below by any column to explore the full dataset.
HUNTINGTIN PROTEIN DOMAIN STRUCTURE
Human HTT (UniProt P42858, 3,142 aa) shown as an example. Hover a segment for details.
13,142 aa
Regions per UniProt P42858; HEAT-repeat solenoid organization per Guo et al. 2018, Nature (cryo-EM structure, PDB 6EZ8).
HUNTINGTIN PROTEIN 3D STRUCTURE
Interactive AlphaFold model of human HTT, residues 1–1,400 of 3,142 (the fragment covering the N-terminal, polyQ, polyP, and HEAT-repeat regions shown above). Drag to rotate, scroll to zoom.
Loading structure…
Structure prediction: AlphaFold DB (Jumper et al. 2021; Varadi et al. 2024), fragment 1 of human HTT (UniProt P42858, residues 1–1400, pLDDT 71.7), mirrored via RCSB PDB computed structure models. Colored by pLDDT confidence. View full entry on AlphaFold DB →
ALL SPECIES
Click a species name for its source database link. The Source column shows which database(s) had this species. Use the checkboxes to select species and copy them as FASTA for tools like Clustal Omega.
Species Name
Common Name
Phylum
Class
Order
Length
PolyQ
PolyP
Sequence
Source
Sources
The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Research, 2025, 53(D1), D609–D617. https://doi.org/10.1093/nar/gkae1010
National Center for Biotechnology Information (NCBI)[Internet]. Bethesda (MD): National Library of Medicine (US), National Center for Biotechnology Information; [1988] – [cited 2026 Jul 26]. Available from: https://www.ncbi.nlm.nih.gov/
How These Orthologs Were Identified
This dashboard combines two independently-run pipelines, one pulling from UniProt and one from NCBI, then merges them into one deduplicated table.
Finding candidates
Neither pipeline runs its own BLAST search from scratch. Both trust each database's own gene/ortholog annotation, unioning two searches to catch what a single search would miss:
UniProt: the UniRef50 cluster containing human HTT, unioned with a direct gene-symbol search. The second catches entries UniRef50's clustering threshold misses.
NCBI: NCBI's own precomputed ortholog group for human HTT (Gene ID 3064), unioned with a direct Entrez search, HTT[gene]. The second is needed because NCBI's ortholog-caller places arthropod HTT in a separate group from human's, a real classification choice rather than a data gap.
Both unioned sets are then filtered against a short blocklist matched against each entry's protein/gene description, dropping known name collisions such as SLC6A4 (the serotonin transporter, sometimes aliased "5-HTT") and a cluster of "huntingtin-like" Lepidoptera genes each database keeps distinct from plain "huntingtin."
The polyQ / polyP rule
A deterministic heuristic, identical across both pipelines: polyQ is the longest run of ≥2 Q starting within the first 25 aa; the repeat tract extends until the first run of 3+ consecutive non-P/Q residues; polyP is the longest run of ≥2 P within that tract; anything left over is "in-between." It's verified on every run against each database's own human reference sequence.
Merging & deduplication
Rows from both sources are grouped by species name (trimmed, case-folded) so a species found in both databases is counted once, not twice. For species found in both:
A curated/reviewed entry (UniProt Swiss-Prot, or NCBI RefSeq Curated) beats a predicted/unreviewed one outright, regardless of length.
Only within the same tier does the longer sequence win.
The non-winning entry isn't discarded. Its accession, link, and polyQ/polyP call are kept alongside the winning row for cross-checking.
Caveats: inclusion is based on the source database's own gene-symbol/ortholog annotation, not independent sequence verification by this project. A low-identity, computationally-predicted match (e.g. an automated Gnomon gene model in a distantly related species) is a much weaker claim than a curated primate entry. Tie-break rules are conventions for picking one row per species, not biological judgments about isoform correctness. UniProt's and NCBI's canonical human reference sequences differ slightly (21 vs. 23 Q's), so "% identity to human HTT" is not directly comparable 1:1 across rows sourced from different databases.