{"data":{"id":"us/37-cfr-appendix-e-to-subpart-g-of-part-1","jurisdiction":"us","citation":"37 CFR Appendix E to Subpart G of Part 1","heading":"Appendix E to Subpart G of Part 1—List of Feature Keys Related to Nucleotide Sequences","body":"Source: World Intellectual Property Organization (WIPO) Handbook on Industrial Property Information and Documentation, Standard ST.25: Standard for the Presentation of Nucleotide and Amino Acid Sequence Listings in Patent Applications (2009).\nKey Description\nallele a related individual or strain contains stable, alternative forms of the same gene, which differs from the presented sequence at this location (and perhaps others).\nattenuator (1) region of DNA at which regulation of termination of transcription occurs, which controls the expression of some bacterial operons; (2) sequence segment located between the promoter and the first structural gene that causes partial termination of transcription.\nC__region constant region of immunoglobulin light and heavy chains, and T-cell receptor alpha, beta, and gamma chains; includes one or more exons depending on the particular chain.\nCAAT__signal CAAT box; part of a conserved sequence located about 75 bp upstream of the start point of eukaryotic transcription units which may be involved in RNA polymerase binding; consensus=GG (C or T) CAATCT.\nCDS coding sequence; sequence of nucleotides that corresponds with the sequence of amino acids in a protein (location includes stop codon); feature includes amino acid conceptual translation.\nconflict independent determinations of the “same” sequence differ at this site or region.\nD-loop displacement loop; a region within mitochondrial DNA in which a short stretch of RNA is paired with one strand of DNA, displacing the original partner DNA strand in this region; also used to describe the displacement of a region of one strand of duplex DNA by a single stranded invader in the reaction catalyzed by RecA protein.\nD-segment diversity segment of immunoglobulin heavy chain, and T-cell receptor beta chain.\nenhancer a cis-acting sequence that increases the utilization of (some) eukaryotic promoters, and can function in either orientation and in any location (upstream or downstream) relative to the promoter.\nexon region of genome that codes for portion of spliced mRNA; may contain 5′UTR, all CDSs, and 3′UTR.\nGC__signal GC box; a conserved GC-rich region located upstream of the start point of eukaryotic transcription units which may occur in multiple copies or in either orientation; consensus=GGGCGG.\ngene region of biological interest identified as a gene and for which a name has been assigned.\niDNA intervening DNA; DNA which is eliminated through any of several kinds of recombination.\nintron a segment of DNA that is transcribed, but removed from within the transcript by splicing together the sequences (exons) on either side of it.\nJ__segment joining segment of immunoglobulin light and heavy chains, and T-cell receptor alpha, beta, and gamma chains.\nLTR long terminal repeat, a sequence directly repeated at both ends of a defined sequence, of the sort typically found in retroviruses.\nmat__peptide mature peptide or protein coding sequence; coding sequence for the mature or final peptide or protein product following post-translational modification; the location does not include the stop codon (unlike the corresponding CDS).\nmisc__binding site in nucleic acid which covalently or non-covalently binds another moiety that cannot be described by any other Binding key (primer__bind or protein__bind).\nmisc__difference feature sequence is different from that presented in the entry and cannot be described by any other Difference key (conflict, unsure, old__sequence, mutation, variation, allele, or modified__base).\nmisc__feature region of biological interest which cannot be described by any other feature key; a new or rare feature.\nmisc__recomb site of any generalized, site-specific or replicative recombination event where there is a breakage and reunion of duplex DNA that cannot be described by other recombination keys (iDNA and virion) or qualifiers of source key (/insertion__seq, /transposon, /proviral).\nmisc__RNA any transcript or RNA product that cannot be defined by other RNA keys (prim__transcript, precursor__RNA, mRNA, 5′clip, 3′clip, 5′UTR, 3′UTR, exon, CDS, sig__peptide, transit__peptide, mat__peptide, intron, polyA__site, rRNA, tRNA, scRNA, and snRNA).\nmisc__signal any region containing a signal controlling or altering gene function or expression that cannot be described by other Signal keys (promoter, CAAT__signal, TATA__signal, -35__signal, -10__signal, GC__signal, RBS, polyA__signal, enhancer, attenuator, terminator, and rep__origin).\nmisc__structure any secondary or tertiary structure or conformation that cannot be described by other Structure keys (stem__loop and D-loop).\nmodified__base the indicated nucleotide is a modified nucleotide and should be substituted for by the indicated molecule (given in the mod__base qualifier value).\nmRNA messenger RNA; includes 5′ untranslated region (5′UTR), coding sequences (CDS, exon) and 3′ untranslated region (3′UTR).\nmutation a related strain has an abrupt, inheritable change in the sequence at this location.\nN__region extra nucleotides inserted between rearranged immunoglobulin segments.\nold__sequence the presented sequence revises a previous version of the sequence at this location.\npolyA__signal recognition region necessary for endonuclease cleavage of an RNA transcript that is followed by polyadenylation; consensus=AATAAA.\npolyA__site site on an RNA transcript to which will be added adenine residues by post-transcriptional polyadenylation.\nprecursor__RNA any RNA species that is not yet the mature RNA product; may include 5′ clipped region (5′clip), 5′ untranslated region (5′UTR), coding sequences (CDS, exon), intervening sequences (intron), 3′ untranslated region (3′UTR), and 3′ clipped region (3′clip).\nprim__transcript primary (initial, unprocessed) transcript; includes 5′ clipped region (5′clip), 5′ untranslated region (5′UTR), coding sequences (CDS, exon), intervening sequences (intron), 3′ untranslated region (3′UTR), and 3′ clipped region (3′clip).\nprimer__bind non-covalent primer binding site for initiation of replication, transcription, or reverse transcription; includes site(s) for synthetic, for example, PCR primer elements.\npromoter region on a DNA molecule involved in RNA polymerase binding to initiate transcription.\nprotein__bind non-covalent protein binding site on nucleic acid.\nRBS ribosome binding site.\nrepeat__region region of genome containing repeating units.\nrepeat__unit single repeat element.\nrep__origin origin of replication; starting site for duplication of nucleic acid to give two identical copies.\nrRNA mature ribosomal RNA; the RNA component of the ribonucleoprotein particle (ribosome) which assembles amino acids into proteins.\nS__region switch region of immunoglobulin heavy chains; involved in the rearrangement of heavy chain DNA leading to the expression of a different immunoglobulin class from the same B-cell.\nsatellite many tandem repeats (identical or related) of a short basic repeating unit; many have a base composition or other property different from the genome average that allows them to be separated from the bulk (main band) genomic DNA.\nscRNA small cytoplasmic RNA; any one of several small cytoplasmic RNA molecules present in the cytoplasm and (sometimes) nucleus of a eukaryote.\nsig__peptide signal peptide coding sequence; coding sequence for an N-terminal domain of a secreted protein; this domain is involved in attaching nascent polypeptide to the membrane; leader sequence.\nsnRNA small nuclear RNA; any one of many small RNA species confined to the nucleus; several of the snRNAs are involved in splicing or other RNA processing reactions.\nsource identifies the biological source of the specified span of the sequence; this key is mandatory; every entry will have, as a minimum, a single source key spanning the entire sequence; more than one source key per sequence is permissible.\nstem__loop hairpin; a double-helical region formed by base-pairing between adjacent (inverted) complementary sequences in a single strand of RNA or DNA.\nSTS Sequence Tagged Site; short, single-copy DNA sequence that characterizes a mapping landmark on the genome and can be detected by PCR; a region of the genome can be mapped by determining the order of a series of STSs.\nTATA__signal TATA box; Goldberg-Hogness box; a conserved AT-rich septamer found about 25 bp before the start point of each eukaryotic RNA polymerase II transcript unit which may be involved in positioning the enzyme for correct initiation; consensus=TATA(A or T)A(A or T).\nterminator sequence of DNA located either at the end of the transcript or adjacent to a promoter region that causes RNA polymerase to terminate transcription; may also be site of binding of repressor protein.\ntransit__peptide transit peptide coding sequence; coding sequence for an N-terminal domain of a nuclear-encoded organellar protein; this domain is involved in post-translational import of the protein into the organelle.\ntRNA mature transfer RNA, a small RNA molecule (75-85 bases long) that mediates the translation of a nucleic acid sequence into an amino acid sequence.\nunsure author is unsure of exact sequence in this region.\nV__region variable region of immunoglobulin light and heavy chains, and T-cell receptor alpha, beta, and gamma chains; codes for the variable amino terminal portion; can be made up from V__segments, D__segments, N__regions, and J__segments.\nV__segment variable segment of immunoglobulin light and heavy chains, and T-cell receptor alpha, beta, and gamma chains; codes for most of the variable region (V__region) and the last few amino acids of the leader peptide.\nvariation a related strain contains stable mutations from the same gene (for example, RFLPs, polymorphisms, etc.) which differ from the presented sequence at this location (and possibly others).\n3′clip 3′-most region of a precursor transcript that is clipped off during processing.\n3′UTR region at the 3′ end of a mature transcript (following the stop codon) that is not translated into a protein.\n5′clip 5′-most region of a precursor transcript that is clipped off during processing.\n5′UTR region at the 5′ end of a mature transcript (preceding the initiation codon) that is not translated into a protein.\n−10__signal pribnow box; a conserved region about 10 bp upstream of the start point of bacterial transcription units which may be involved in binding RNA polymerase; consensus=TAtAaT.\n−35__signal a conserved hexamer about 35 bp upstream of the start point of bacterial transcription units; consensus=TTGACa [ ] or TGTTGACA [ ].","path":["Title 37—Patents, Trademarks, and Copyrights","CHAPTER I—UNITED STATES PATENT AND TRADEMARK OFFICE, DEPARTMENT OF COMMERCE","SUBCHAPTER A—GENERAL","PART 1—RULES OF PRACTICE IN PATENT CASES","Subpart G—Biotechnology Invention Disclosures"],"source_url":"https://www.ecfr.gov/api/versioner/v1/full/2026-08-25/title-37.xml","current_through":"2026-08-25","vintage":"","retrieved_at":"2026-08-27T02:25:45Z","sha256":"7c3f561e6dfc1411b95e79ff00ea83141d14fdc4a833fef67f4dfc26cbda0165","source_id":"us-cfr","stale":true,"prev":"us/37-cfr-appendix-d-to-subpart-g-of-part-1","next":"us/37-cfr-appendix-f-to-subpart-g-of-part-1"},"notice":"GroundRules: Original legal text. Not legal advice."}
