Molecular Modeling Database (MMDB) Help Document
发布时间:2026-09-21 | 浏览:1
This help document provides detailed descriptions of the Entrez Structure database content, search system, and display formats. The " How To " page provides quick start guides for some common types of searches. Once records of interest are retrieved, follow Entrez's "Links" to discover associations among previously disparate data . The Entrez Help document provides additional information about the search system and the databases it can be used to search.
What are macromolecular structures ?
Four levels of protein structure (primary, secondary, tertiary, quaternary)
Experimental methods (X-ray crystallography, NMR)
How can 3D structures be used to learn more about proteins and other biomolecules?
identify representative 3D structures for protein families
examine sequence-structure-function relationships ( illustrated example )
view 3D structures of conserved core motifs
identify putative active site residues
Useful Features of the Molecular Modeling Database
Facilitate computation on 3D structure data
Analysis of individual structures and relationships among them
biological and geometrical features within 3D structures
conserved protein domain annotations
evolutionary relationships among 3D structures
functional relationships among 3D structures
Interactive views of sequence-structure relationships
Connections between 3D structure records and associated literature, molecular, and chemical data
Content of the Molecular Modeling Database
Source database
RSCB Protein Data Bank (PDB)
How are the data processed at NCBI?
content validation
deposit sequence and chemical data into Entrez Protein, Nucleotide, and PubChem databases
identify biological units (oligomeric states) ( illustrated example )
author/software determination
apply transformations derived from crytallographic symmetry
compare biological units within a record to each other to identify distinct forms
note about biological units in merged PDB split files
technical note about asymmetric unit
merge PDB split files
illustrated example : viral capsid
illustrated example : rat liver vault
illustrated example : ribosome
identify interactions among molecular components
4 Å interatomic distance
5 or more contacts
rank interactions
identify geometrical features
secondary structures in protein molecules
3D domains in protein molecules
identify the gene that corresponds to each protein
identify relationships among 3D structures
find similar 3D structures using VAST algorithm
create links to associated data throughout the Entrez system
Record types ( illustrated examples )
experimental methods (X-ray crystallography, NMR, other)
molecule types (protein, DNA, RNA)
Update frequency
INPUT: Search Tips
Allowable search terms
text terms (names of proteins, bound chemicals, authors, etc.)
unique identifiers
database subset
Basic search (& search details )
Advanced search ( Search builder , Show index list , History )
Complex Boolean query
Range search (range of dates, molecular weights, etc.)
complete list of search field names, abbreviations, and descriptions
tips about search field abbreviations , use of quotes around query terms , and use of wild-card (*)
Link from other Entrez database ( illustrated example )
traverse from sequence/literature/small molecule/other databases to 3D structures
links from protein sequence records to 3D structures
OUTPUT: Search Results
Document summary (docsum) page : list of records found ( illustrated example )
"Display Settings" menu
Filter your results
Refine your results
Find related data
Similar Structures
PubMed Central Full Text
PubMed Citations
Conserved Domain Family
Conserved Domain Superamily
Conserved Domains
PubChem Compound
PubChem Substance
Related Protein
View details for an individual 3D structure record
Structure Summary Page : What information is displayed for each macromolecular structure? ( illustrated example )
Record identifiers
Descriptive information
PDB deposit date
MMDB update date
Source organism
Similar structures: VAST+
Experimental method
Display options
Default biological unit
All biologicial units
Asymmetric unit
Structure images
Molecular graphic
Interactions schematic
Download structure data (save 3D structure record)
Single 3D structure
All 3D structures
PDB source file
Additional details about structure data download options
annotated illustration of download options
details about data saved in each file format
save image of 3D structure
save structure components
Molecular components
Tabular list of molecular components
Column headers: label, count, molecule
Molecule label, count, & name
Protein annotation graphic
Domain families (protein classification)
Molecule label, count, & name
Thumbnail graphic
Molecule label, count, & name
Thumbnail graphic
Non-standard Biopolymers
Molecule label, count, & name
Web API : URL format for displaying or saving a structure record
parameters & allowable values
examples of URLs for displaying or saving 3D structure records
Citing the Molecular Modeling Database
Additional references
Four Levels of Protein Structure
Experimental Methods
How can 3D structures be used to learn more about proteins and other biomolecules?
Facilitate computation on 3D structure data
Find structures for a gene/protein product of interest or its homologs.
Find 3D structures bound to a specific chemical (e.g., aspirin).
Align a query protein to a similar sequence from a 3D structure and interactively view sequence/structure relationships.
Identify structures within the database that are similar to each other, regardless of their degree of sequence similarity.
Analysis of individual structures and relationships among them
Interactive views of sequence-structure relationships
Connections between 3D structure records and associated literature, molecular, and chemical data
Source Database
How are the data processed at NCBI?
Content Validation :
When PDB structure records are imported into MMDB, the information in each structure record is reorganized and validated in a way that enables cross-referencing between the chemistry and the three-dimensional structure of macromolecules. While the PDB data model provides an elegant and concise description of a crystal structure, there is no one-to-one correspondence between a site, a structure, and an atom in the chemical sense. MMDB provides this chemical information in an explicit manner. Its data specification includes a description of a biopolymer's spatial structure, a description of how it is organized chemically, and a set of pointers linking the two.
The first step in creating MMDB is getting an accurate sequence that is consistent with the atom site coordinates in PDB. For example:
The SEQRES records in an original PDB file are generally intended to represent the molecule that was purified, crystallized, and measured. However, it might not have been possible to experimentally resolve the atomic coordinates for all of the amino acids in some structures, especially in flexible regions of proteins such as N- and C- terminals. In addition, sometimes the atomic coordinates might indicate the presence of additional residues not listed in the SEQRES records. In the latter case, MMDB derives the biopolymer sequence from the atomic coordinates and not from the original SEQRES records. The derived biopolymer sequence will then appear in the MMDB record, and in the SEQRES records of the PDB-formatted file saved from the MMDB database.
Some PDB records may have discontinous residue numbers, which exist in a free text field. MMDB assigns a consecutive series of positive integers to residues in biopolymers, using a numerical data field. This ensures correspondence between the residue numbers in the structure file and those in the corresponding protein and/or nucleotide sequence records.
The second step is to construct a complete chemical graph for the molecule, representing all bonds and chirality. An important component of this second step matches the amino acid and nucleotide groups defined by PDB against a dictionary that defines all bond and atom types.
The third and final step is to recover disorder information in the structure.
(Note: Because such changes may occur during data processing, the content of a PDB-formatted file that you save from the MMDB database might differ from the original PDB file. )
Deposit sequence and chemical data into Entrez Protein , Nucleotide , and PubChem databases :
In addition to providing the spatial (x,y,z) coordinates of every atom in a 3D macromolecular structure, a structure record includes the sequence data for each component nucleotide (DNA, RNA) and/or protein molecule. As part of MMDB data processing , the sequence data for each molecule are deposited into the Entrez Nucleotide or Entrez Protein database, as appropriate. The data processing procedures for those databases, in turn, identify relationships (i.e., similarities) among the sequence data from 3D structures and the other sequences in those databases, facilitating the use of 3D structure data to learn more about proteins and other biomolecules . A structure record may also include bound chemicals. Data records for those chemicals are deposited into the PubChem Substance database, and then linked to corresponding records in the non-redundant, curated PubChem Compound database. This makes it possible, for example, to find 3D protein structures bound to a specific chemical (e.g., aspirin) , even if submitters of 3D structures used various names or abbreviations for a given chemical.
Identify biological units (oligomeric states) :
What is a biological unit? The biochemically active form of a biomolecule can range from a monomer (single protein molecule) to an oligomer of 100+ protein molecules , and is referred to as " biological unit " for brevity. The raw data present structure records resolved by x-ray crystallography or neutron diffraction of a crystal are often casually referred to as the " asymmetric unit ." These data can represent either: (a) the complete biological unit, (b) a portion of the biological unit, or (c) multiple copies of the biological unit, as in the human hemoglobin examples shown below. Authors of structure records use programs such as PISA to identify the biological unit within a structure record. If multiple interpretations of the biological unit exist, the author may choose to annotate the various interpretations in their record. The MMDB data processing pipeline applies several procedures to identify a structure's biological unit(s) and displays it by default on a structure summary page. (See technical note about asymmetric unit.) The asymmetric unit is equivalent to the biological unit in approximately 60% of structure records resolved by x-ray crystallography or neutron diffraction of crystals. In the remaining 40% of the records, the asymmetric unit represents a portion of the biological unit that can be reconstructed using crystallographic symmetry , or it represents multiple copies of the biological unit. Additionally, some structures exceed the size limits implicit to the PDB file format and are therefore split by PDB into several files. In those cases, the biological unit might be spread across multiple PDB files. The MMDB data processing pipeline merges the split files into a single structure record. In such cases, " asymmetric unit " is the only display option for merged PDB split files from crystallographic studies , because the biological unit of the complete structure is not specified in a computer readable way in the PDB source files. The structure summary page for a merged crystallographic structure therefore simply uses the label of " asymmetric unit" above the molecular graphic, because it represents the unification of raw data from the original PDB files. The asymmetric unit can represent the structure's complete biological unit, a portion of the biological unit, or multiple copies of the biological unit. In the case of structures resolved by electron microscopy (EM) or nuclear magnetic resonance (NMR) , the term "asymmetric unit" does not apply, and the term "biological unit" is shown instead on the summary page for a merged structure from either of those technologies. Please refer to the corresponding publication for a structure, if/as available, for the author's description of its biologically active form.
Asymmetric unit (raw data) → Biological unit (default display) Example: -- As an example of the varying degrees to which a biological unit can be represented by the raw data in a structure record, compare the following records for human hemoglobin . Each one contains the spatial coordinates and sequence data for a different number of protein molecules, yet the fundamental biological unit in all three structures is a tetramer consisting of two alpha, two beta subunits, and four heme groups. By default, an MMDB structure summary page displays the biological unit:
Procedures to identify the biological unit(s) within a structure record :
Asymmetric unit (technical note):
The raw data in a structure record (generated by x-ray crystallography or neutron diffraction) are often casually referred to as the " asymmetric unit ." These data, which were submitted by the author and stored in the source PDB record , can represent either : (a) the complete biological unit (i.e, the biochemically active form of a biomolecule); (b) a portion of the biological unit; or (c) multiple copies of the biological unit, as shown in the illustrated example of three different human hemoglobin structure records. The display options on an MMDB summary page for an individual structure allow you to view your choice of biological unit(s) or asymmetric unit, with the biological unit shown by default. The "asymmetric unit" is equivalent to the biological unit in approximately 60% of structure records . The concepts of asymmetric unit and biological unit do not apply to structure records resolved by experimental methods other than x-ray crystallography and neutron diffraction. Note: The technical definition of asymmetric unit is somewhat different from its casual meaning. Technically, an asymmetric unit is the smallest part of a 3D structure from which the complete structure can be built using a specific set of rotational and translational matrices that describe the symmetry of the structure.
Merging PDB split files into a single MMDB structure record
Some structures exceed the size limits implicit to the PDB file format and are therefore split into several PDB files. The MMDB data processing procedures merge the PDB split files into a single structure record. The merged structures now make it possible to display and/or download large macromolecular structures in their entirety, and to interactively view the sequence-structure relationships using either iCn3D , a web-based 3D viewer that loads the structure within the web page without the need to install a separate application, or the stand-alone Cn3D 4.3 ( install ). Please note that " asymmetric unit " is the only display option for merged PDB split files from crystallographic studies , because the biological unit of the complete structure is not specified in a computer readable way in the PDB source files. The structure summary page for a merged crystallographic structure therefore simply uses the label of " asymmetric unit" above the molecular graphic, because it represents the unification of raw data from the original PDB files. The asymmetric unit can represent the structure's complete biological unit, a portion of the biological unit, or multiple copies of the biological unit. In the case of structures resolved by electron microscopy (EM) or nuclear magnetic resonance (NMR) , the term "asymmetric unit" does not apply, and the term "biological unit" is shown instead on the summary page for a merged structure from either of those technologies. Please refer to the corresponding publication for a structure, if/as available, for the author's description of its biologically active form. Examples of merged structures, illustrated below, include the:
viral capsid by Xie et al.
rat liver vault by Tanaka et al.
ribosome structure by Nobel Laureate V. Ramakrishnan
You can also retrieve all merged files from the Molecular Modeling Database, if desired.
Example: The viral capsid for the Adeno-associated Virus Serotype 6 (Aav-6) by Xie et al. was split into PDB records 1VU0 , 1VU1 , 3TSX , and was merged at MMDB into a single record with the MMDB ID 99554 :
Example: The rat liver vault by Tanaka et al. was split into PDB records 2ZUO , 2ZV4 , 2ZV5 , and was merged at MMDB into a single record with the MMDB ID 99596 : (Note: The merged file represents half of the biological unit , as it was submitted by the author. The procedures to identify biological units cannot be applied in an automated way to a merged file ; therefore, the asymmetric unit is diplayed instead. Please refer to the corresponding publication for a structure for the author's description of the biologically active form. )
Example: The ribosome structure by Selmer, Dunham, Murphy, Weixlbaumer, Petry, Kelley, Weir, and Ramakrishnan , the 2009 Nobel Laureate in Chemistry , was split into PDB records 2XFZ , 2XG0 , 2XG1 , 2XG2 , and was merged at MMDB into a single record with the MMDB ID 99580 : (Note: The merged file represents two copies of the biological unit , as submitted by the author. The procedures to identify biological units cannot be applied in an automated way to a merged file ; therefore, the asymmetric unit is diplayed instead. Please refer to the corresponding publication for a structure for the author's description of the biologically active form. )
In summary, the merged structure files, such as the viral capsid , the rat liver vault , and the ribosome illustrated above, now make it possible to view and/or download large macromolecular structures in their entirety, and to interactively view the sequence-structure relationships using the free stand-alone Cn3D 4.3 program ( install ), or the free web-based iCn3D viewer. You can also retrieve all merged files from the Molecular Modeling Database, if desired. Please refer to the corresponding publications for those structures, if/as available, for the author's description of their biologically active form.
Identify interactions among molecular components :
As part of MMDB data processing , the spatial coordinates in a structure record are analyzed to identify interactions among the structure's molecular components . Interactions are reported on an MMDB Summary Page as an interactions schematic if they meet the following thresholds:
Identify geometrical features :
Identify the gene that corresponds to each protein :
Identify relationships among 3D structures :
Create links to associated data throughout the Entrez system:
As noted in the page on discovering associations among previously disparate data , the Entrez retrieval system is designed to provide integrated access to previously disparate data and make it possible to collect related information on a topic of interest within and across Entrez databases. MMDB therefore identifies such associations during data processing and presents them as " Related Information " menus on search results pages. Many of the links are also available on individual structure records . There are two broad categories of Links:
Update Frequency
Allowable search terms
suppressor OR inhibitor
NF1 OR neurofibromin OR neurofibromatosis
PTGS1 OR "prostaglandin endoperoxide synthase 1" (see note about use of quotes )
Basic search (& search details )
Advanced search ( Search builder , Show index list , History )
Complex Boolean query
Range search (range of values in numerical fields such as dates, counts, and resolution ).
Just enter search terms without specifying search fields, other limits, or Boolean operators.
The " Search Details " box in the right margin of the search results page shows exactly how Entrez parsed and handled your query. If desired, you can edit the query in that box and press the "Search" button to run the modified query. The " See more... " link a the bottom of the "Search Details" box opens a more detailed display:
The Query Translation box shows the search strategy used to run the search
To edit the search in the Query Translation box, add or delete terms and then click Search.
Click URL to display the current search as a URL to bookmark for future use. Searches created using History numbers can not be saved using the URL feature.
You may also save your search using My NCBI .
The Result number link retrieves the documents found and displays them in a search results page.
Translations details how each term was translated using Entrez's search rules and syntax for the database.
User Query shows the search terms as you entered them in the search box and any syntax errors with the query.
The Limits page allows you to restrict your search in various ways. At a minimum, the Limits page displays the list of available search fields . You can do a separate search for each term or phrase in your query, as shown in sample Search #2 and #3 to the right, and select the desired search field for each one. (If desired, you can then combine the searches by using the Search Builder or History section of the Advanced Search page.) For some databases, the Limits page also provides other commonly used options, as check boxes and/or pull-down menus, for restricting your search results to records with specific characteristics . These check boxes and pull-down menus generally represent a commonly used subset of the choices that are available from the Advanced Search page and are placed on the Limits page for easy access. IMPORTANT NOTE : Once you have used a particular Limit, warning sign will appear near the top of your search results page that indicates which Limit(s) are currently in effect, for example: Note that the Limit will remain in effect for all subsequent searches in the current database unless you change or remove that limit . In the illustrated example above, any search you do will be limited to the Titles of records, until you remove the limit.
At a minimum, the Limits page displays the list of available search fields . You can do a separate search for each term or phrase in your query, as shown in sample Search #2 and #3 to the right, and select the desired search field for each one. (If desired, you can then combine the searches by using the Search Builder or History section of the Advanced Search page.)
For some databases, the Limits page also provides other commonly used options, as check boxes and/or pull-down menus, for restricting your search results to records with specific characteristics . These check boxes and pull-down menus generally represent a commonly used subset of the choices that are available from the Advanced Search page and are placed on the Limits page for easy access.
IMPORTANT NOTE : Once you have used a particular Limit, warning sign will appear near the top of your search results page that indicates which Limit(s) are currently in effect, for example: Note that the Limit will remain in effect for all subsequent searches in the current database unless you change or remove that limit . In the illustrated example above, any search you do will be limited to the Titles of records, until you remove the limit.
Build a search one step at a time .
Browse the index of any search field and add term(s) of interest from the index to the active query box at the top of the page.
View your search History and combine or subtract searches from each other.
Complex Boolean query
The "Search Builder" section of the Advanced Search page allows you to build your query step by step, adding a new search term and selecting a new search field at each step. It also allows you to browse the index of any search field to view the available terms.
To build a query : (1) Select the Search Field of interest using the pull-down menu. (2) Type a term(s) in the text box beside the search field menu. Or , use the " Show index list " link to see the index of the search field and select the desired term from the index. ( tips on using the "Show Index List" ) (3) Select the Boolean operator (AND, NOT, OR) that should precede the term when it is added to the active query at the top of the page. Continue the above steps, as desired, to add more term/search field combinations to your query.
As you use the Search Builder, the grey text box at the top of the page will show your current query . You can manually edit the current query by clicking the "Edit" link beneath the grey text box. That will allow you to type terms/search numbers/etc. directly into the box, add parentheses for nesting if desired, change Boolean operators, etc. Press the Search button to display the records retrieved by your search (i.e., it displays the search results page). Click on the " Add to history " link if you prefer to simply add the query to your search history and remain on the Advanced Search page, where you can continue building your query.
Tips on using the " Show Index List " function on the Advanced Search page: The " Show Index List " function allows you to browse the index of any Search Field . If you select a search field and press the "Show Index" link without entering a term in the box, you will be taken to the top of the index. If you enter a term first , you will be taken to the part of the index that contains your term (or the closest alphabetical location, if your term is not present in the index). The number of records that contain the term will appear in parentheses. You can also browse the index to explore the variety of terms available (for example, select "All Fields", enter "Huntington", and click on the "Show Index" link to see additional spellings and/or related terms, such as Huntington disease, Huntington's, Huntington's disease). To select a range of terms from the index, use the Shift key while selecting the first and last term. Then use the AND, OR, or NOT buttons to add that group of terms to the active query. To select multiple terms that do not fall within a continuous range from the index, use the Control key while selecting the terms of interest. Then use the AND, OR, or NOT buttons to add that group of terms to the active query. Note: When multiple terms are selected from the index window, they are OR'ed together within parentheses and then appended to your query with whatever Boolean operator you have selected.
The "History" section of the Advanced Search page displays the searches you have done in the current database.
You can combine or subtract searches from each other by entering the search numbers and the AND, OR, or NOT Boolean operators in the query box, for example: #2 AND #3 . If the query contains several search numbers and Boolean operators, the Boolean operators are processed from left to right unless parentheses are used for nesting. If parentheses are used, the portions of the query in parentheses will be processed first, then the remaining Boolean operators will be processed from left to right.
Additional details about Search History:
The Search History will be lost after 8 hours of inactivity . (To save a search indefinitely, click on the search # and select " Save in My NCBI .)
Click "Clear History" to delete all searches from History.
Entrez will move a search statement number to the top of the History if a new search is the same as a previous search.
History search numbers may not be continuous because some numbers are assigned to intermediate processes, such as displaying a citation in another format.
The maximum number of searches held in History is 100. Once the maximum number is reached, PubMed will remove the oldest search from the History to add the most current search.
A separate Search History will be kept for each database, although the search statement numbers will be assigned sequentially for all databases.
PubMed uses cookies to keep a history of your searches. For you to use this feature, your Web browser must be set to accept cookies .
Database records that you have copied to the Clipboard are represented by the search number #0, which may be used in Boolean search statements. For example, to limit the records you have collected in the Clipboard to those from human, use the following search: #0 AND human[organism]. This does not change or replace the Clipboard contents.
Enter a search in command language , specifying your exact combination of desired search terms, search fields, and Boolean operators, as shown in the examples to the right. The syntax is: term[field] BOOLEAN term[field] BOOLEAN term[field] etc.
Search Field names must be placed in square brackets [] , and can be written as either the full name, for example, [Database], or as the corresponding search field abbreviation , for example, [db] ( additional examples ).
Boolean operators ( AND , OR , NOT ) must be written in UPPER CASE .
Boolean operators are processed from left to right unless parentheses are used for nesting. If parentheses are used, the portions of the query in parentheses will be processed first, then the remaining Boolean operators will be processed from left to right.
Boolean operators can also be used to combine or subtract searches from each other (i.e., to find the union, difference, or intersection of the data sets retrieved by various searches). To do this, use the Search History section of the Advanced Search page and simply enter the search numbers and desired Boolean operators in the query box. For example, to identify the records that were retrieved by Search #2 of your search history, and also by Search #3, you could enter the following query: #2 AND #3 To identify the records that were retrieved by Search #2 but not by Search #3, you could enter the following query: #2 NOT #3
Range queries are constructed by specifying a lower and upper numerical value separated by a colon (:) to specify the range, followed by a search field name or abbreviation in square brackets, as shown in the examples to the right. You can insert a space on each side of the colon but that is not necessary; the search will work either way. All dates and all ' counts ' (such as residue counts, molecule counts, etc.) fields can be range queried. Apart from that, there are two additional fields that can be range queried: Resolution [RESO] in the Entrez Structure database, and MolWeight [MWT] in the Entrez Protein database (from which you can link to the Structure database). Range queries on Resolutions [RESO] (in angstroms) must have the following format: fromResolution : toResolution [RESO] Range queries on MolecularWeights [MWT] (in daltons) must have the following format: fromMolecularWeight : toMolecularWeight [MWT] Note that searches by molecular weight are currently possible only in the Entrez Protein database. When you are searching that database, simply append "AND srcdb_pdb[prop]" to your query if you want to retrieve only the protein sequences that were derived from 3D structure records. For example: _____:_____[molwt] AND srcdb_pdb[prop] That will retrieve protein sequences that fall within the specified molecular weight range and that were derived from Protein Data Bank (PDB), the source database for 3D structure records. A specific example is provided in Search #10 to the right. Range queries on Dates have a similar format: FromDate : ToDate [fieldname] Note: The FromDate and ToDate values can specify an exact date, a month, or a year, and are written in the format: YYYY/MM/DD, YYYY/MM, or YYYY. The search fields summary table includes the names and abbreviations for the various "date" fields. Range queries on " counts " have the format: FromCount : ToCount [fieldname] Note: The FromCount and ToCount values are integers. The search fields summary table includes the names and abbreviations for the various "counts" fields.
Range queries on Resolutions [RESO] (in angstroms) must have the following format: fromResolution : toResolution [RESO] Range queries on MolecularWeights [MWT] (in daltons) must have the following format: fromMolecularWeight : toMolecularWeight [MWT] Note that searches by molecular weight are currently possible only in the Entrez Protein database. When you are searching that database, simply append "AND srcdb_pdb[prop]" to your query if you want to retrieve only the protein sequences that were derived from 3D structure records. For example: _____:_____[molwt] AND srcdb_pdb[prop] That will retrieve protein sequences that fall within the specified molecular weight range and that were derived from Protein Data Bank (PDB), the source database for 3D structure records. A specific example is provided in Search #10 to the right. Range queries on Dates have a similar format: FromDate : ToDate [fieldname] Note: The FromDate and ToDate values can specify an exact date, a month, or a year, and are written in the format: YYYY/MM/DD, YYYY/MM, or YYYY. The search fields summary table includes the names and abbreviations for the various "date" fields. Range queries on " counts " have the format: FromCount : ToCount [fieldname] Note: The FromCount and ToCount values are integers. The search fields summary table includes the names and abbreviations for the various "counts" fields.
Range queries on MolecularWeights [MWT] (in daltons) must have the following format: fromMolecularWeight : toMolecularWeight [MWT] Note that searches by molecular weight are currently possible only in the Entrez Protein database. When you are searching that database, simply append "AND srcdb_pdb[prop]" to your query if you want to retrieve only the protein sequences that were derived from 3D structure records. For example: _____:_____[molwt] AND srcdb_pdb[prop] That will retrieve protein sequences that fall within the specified molecular weight range and that were derived from Protein Data Bank (PDB), the source database for 3D structure records. A specific example is provided in Search #10 to the right. Range queries on Dates have a similar format: FromDate : ToDate [fieldname] Note: The FromDate and ToDate values can specify an exact date, a month, or a year, and are written in the format: YYYY/MM/DD, YYYY/MM, or YYYY. The search fields summary table includes the names and abbreviations for the various "date" fields. Range queries on " counts " have the format: FromCount : ToCount [fieldname] Note: The FromCount and ToCount values are integers. The search fields summary table includes the names and abbreviations for the various "counts" fields.
Range queries on Dates have a similar format: FromDate : ToDate [fieldname] Note: The FromDate and ToDate values can specify an exact date, a month, or a year, and are written in the format: YYYY/MM/DD, YYYY/MM, or YYYY. The search fields summary table includes the names and abbreviations for the various "date" fields. Range queries on " counts " have the format: FromCount : ToCount [fieldname] Note: The FromCount and ToCount values are integers. The search fields summary table includes the names and abbreviations for the various "counts" fields.
Range queries on " counts " have the format: FromCount : ToCount [fieldname] Note: The FromCount and ToCount values are integers. The search fields summary table includes the names and abbreviations for the various "counts" fields.
Link from other Entrez Database
Structure - Protein sequence records that have a direct association with the structure record because at least one of the following is true: (a) the protein sequence record was derived directly from a 3D structure record (as described in MMDB data processing ); (b) the accession number of the protein sequence record was listed in the DBREF record of the PDB source file; (c) the protein accession listed in the DBREF record of the PDB source file is also found in an Entrez Gene record, and that Gene record also has links to other protein accession(s); in such a case, all of the protein accessions in the Entrez Gene record will have "Structure" links (and will show a thumbnail image of a corresponding 3D structure in their protein sequence record display); or (d) the protein is identical in composition and sequence length to any of the proteins noted in (a), (b), or (c).
Document Summary (DocSum) page
The initial search results provide a list ( document summary , or " docsum ") of the structure records that contain your search term , which can appear in any field of the record , unless a search field was specified in the query. If desired, you can narrow your search by restricting the query to a search field of interest or adding more terms with a Boolean AND. Alternatively, you can broaden your search by adding more terms (e.g., synonyms) to your query with a Boolean OR . Once you are satisfied with your search results, click on the thumbnail image, PDB Accession , or MMDB ID of any record on the DocSum page to view its structure summary page . In addition, the following options are available for viewing the search results:
Advanced Search
Display Settings
" Send To " menu
Subsets of Results
Filter your results
Refine your results
Find Related Data
Similar Structures
See My NCBI help for:
View details for individual structure record
Illustrated example
Display settings
Summary -- a summary of all of the structure records (default) retrieved by your search, or for those you have selected with checkboxes, in HTML format . The information shown for each record may include the following, as available: Title (description) of the structure record Enzyme Commission (EC) number , if available Taxonomy (source organism(s)) of the protein and/or nucleotide sequences that comprise the structure Molecular components (proteins, nucleic acids, chemicals) present the structure Modification date MMDB ID and PDB ID A subset of links to additional information about the structure, including a " View in iCn3D " link that opens an interactive view of the 3D structure in NCBI's free web-based structure viewing program, as well as links to related data in other Entrez databases. (Note: The " Find Related Data " menu in the right margin of the search results page provides a complete list of links. That menu retrieves related data for all records (default) retrieved by your search, or for the subset of records you have selected with checkboxes.)
Title (description) of the structure record
Enzyme Commission (EC) number , if available
Taxonomy (source organism(s)) of the protein and/or nucleotide sequences that comprise the structure
Molecular components (proteins, nucleic acids, chemicals) present the structure
Modification date
MMDB ID and PDB ID
A subset of links to additional information about the structure, including a " View in iCn3D " link that opens an interactive view of the 3D structure in NCBI's free web-based structure viewing program, as well as links to related data in other Entrez databases. (Note: The " Find Related Data " menu in the right margin of the search results page provides a complete list of links. That menu retrieves related data for all records (default) retrieved by your search, or for the subset of records you have selected with checkboxes.)
Summary (text) -- a summary of the records retrieved by your search, in plain text format . By default, all records from your search result are listed. If you are interested only in specific records, select their checkboxes, select the desired display settings, and press "Apply" to view only those records. The information shown for each record is the same as in the " Summary " format described above, but does not include the subset of links to additional information.
UI List -- a list of the unique identifiers (UI's) for all of the structure records (default) retrieved by your search, or for those you have selected with checkboxes.
By default, 20 documents are listed per page. If desired, decrease (to a minimum of 5) or increase (to a maximum of 200) the number of documents displayed per page then press the "Apply" button.
Search results are displayed in order of decreasing relevance with respect to the query. Many search fields have a score or rank associated with them; for example, the Title and Organism fields have a high rank, while the PdbComment field has a lower rank. The presence of a search term in any one or more of the fields is scored accordingly by the search system, and the total score given to a hit is used in determining its relevance to the query and therefore its placement on the search results page.
Additional options are available to sort records by descending or ascending order of PDB Accession , PDB Deposit Date , MMDB Entry Date , Protein Molecule Count , DNA Molecule Count , RNA Molecule Count , and Chemical Count .
Technical note : If you retrieve all records in the database by searching the Filter field for All[Filt] , the records are simply displayed in descending order of UID (i.e., MMDB ID).
Saves all the hits retrieved by your search into a plain text file, in either "Summary (text)" or "UI List" format .
Copies all the hits retrieved by your search (default), or those you have selected with check boxes, into a Clipboard , which temporarily stores up to 500 items (they will be lost after 8 hours of inactivity). Click on the "Clipboard: XX items" link in the upper right corner of the page to view the items in any format for up to 8 hours after your last activity in the database. The Clipboard will not add an item that is currently in the Clipboard; it will not create duplicate entries. You can remove items from the Clipboard, if desired. Entrez uses cookies to add your selections to the Clipboard. For you to use this feature, your Web browser must be set to accept cookies. Items in the Clipboard are represented by the search number #0, which may be used in Boolean search statements. For example, to limit the items you have collected in the Clipboard to those from human, use the following search: #0 AND human[organism]. This does not affect or replace the Clipboard contents. The Clipboard's " Send to " menu offers you the same "File" and " Collections " options as offered on the original search results page. The latter option saves all items (default), or the subset of items selected with check boxes, indefinitely in the My NCBI Collections section of your My NCBI account.
Saves all the hits retrieved by your search (default), or those you have selected by using their checkboxes, into the My NCBI Collections section of your My NCBI account.
Filter your results
Refine your results
Protein Domain Families - subsets of 3D structures in your search results that contain protein molecules annotated with conserved domains , inferring protein function:
Families - 3D structures containing at least one protein molecule annotated with a specific hit to a conserved domain , suggesting a high confidence level for the inferred function of the protein. Subsets under this header list the top five conserved domains found as specific hits in the structures retrieved by your search. The number in parentheses represents the subset of structures from your search results that contain one or more protein molecules annotated with a specific hit to the listed domain; clicking on the number will retrieve that subset of structure records. The " All XX Families " link will open a list of all the conserved domain models (in the Conserved Domain Database ) that had at least one specific hit to a protein component of any structure found by your search.
Superfamilies - 3D structures containing at least one protein molecule annotated with a any type of hit to a conserved domain , inferring that protein's function and therefore its membership in a superfamily . Subsets under this header show the top five conserved domain superfamilies found in the structures retrieved by your search. The number in parentheses represents the subset of structures from your search results that contain one or more protein molecules annotated with the listed superfamily; clicking on the number will retrieve that subset of structure records. The " All XX CDD Superfamilies " link will open a list of all the superfamilies (in the Conserved Domain Database ) that were annotated on proteins components of the structures found by your search.
Complexes - subsets of 3D structures in your search results that contain specific combinations of molecular components :
Protein-Protein - 3D structures that contain at least two protein molecules.
Protein-DNA - 3D structures that contain at least one protein molecule and one DNA molecule.
Protein-RNA - 3D structures that contain at least one protein molecule and one RNA molecule.
Protein-Chemical - 3D structures that contain at least one protein molecule and one chemical.
Literature - subsets of 3D structures in your search results that contain links to published literature:
PubMed - 3D structures that contain links to bibliographic information (article title, authors, abstract, journal name, etc.) in the PubMed database, for publications that describe the structures.
PMC - 3D structures that contain links to full text articles that describe the structures in the PubMed Central (PMC) free digital archive of biomedical and life sciences journal literature.
Taxonomy - 3D structures that contain at least one molecular component (protein or nucleotide sequence) from the various organisms that were found in the search results.
Subsets under this header show the five organisms most frequently found in the structures retrieved by your search. The number in parentheses represents the subset of structures from your search results that contain one or more protein molecules from the listed organism; clicking on the number will retrieve that subset of structure records. The " All XX Organisms " link shows the total number of different organisms found in the search results and opens that list of organisms in the NCBI Taxonomy database.
Find related data
When a 3D structure includes one or more protein molecules, it also includes sequence data for each protein molecule.
In addition to that sequence data, the structure record may also include a cross-reference other protein sequences (often Swiss-Prot) by listing their accession number in the "DBREF" record of the PDB source file. (The documentation about PDB file format provides more information about the various "records" (data fields) that are present in PDB source files.)
If the protein accession from the "DBREF" record is also listed in an Entrez Gene record, a link is created between the structure record and the Gene record.
View details for an individual structure record
Search box & button
Record Identifiers
Descriptive information
PDB deposit date
MMDB update date Source organism Similar structures: VAST+ Experimental method
Source organism
Similar structures: VAST+
Experimental method
Display Options
Default Biological Unit
All Biological Units
Asymmetric Unit
Biological Unit N
Determined by...
Structure Images
Molecular Graphic
Interactions Schematic
Download Structure Data
Format ASN.1 (Cn3D) PDB XML JSON PNG (image)
Data Set Single 3D Structure All 3D Structures Alpha Carbons PDB source file
Additional details annotated illustration of download options scope of data saved save image save structure components
Molecular Components
Proteins Gene symbol 3D Domains Domain Families (Protein classification) - Specific Hits - Superfamilies - Multidomains
Search box & button
Structure Record Identifiers
Descriptive Information
Display Options
Biological Unit N
provides a type classification based on a comparison of the biological units identified in the structure record, if the record contains multiple biological units. If two or more biological units meet a threshold for sequence and structural similarity, they will receive the same type code; if they do not meet that threshold, they are considered distinct from each other and received different type codes.
indicates the oligomeric state (dimer, trimer, tetramer, etc.) and the method by which it was determined
presents a schematic diagram of interactions , a molecular graphic , options to download the structure data , and a table of molecular components .
Structure Images
STATIC IMAGE: Upon first opening a structure summary page , the molecular graphic shows a static image of the 3D structure . The static image generally shows the default biological unit of the structure.
Click the button in the lower left corner of the static image to load an interactive view that uses iCn3D ("I see in 3D"), NCBI's web-based 3D structure viewer.
The interactive display will load only if your web browser supports WebGL . If it doesn't, the static image will be shown instead. To see the interactive view, modify the settings in your web browser to enable WebGL, or, if needed, update your web browser to a newer version that supports WebGL. (See the WebGL site for more information about compatibility with various web browsers.)
The actions that happen when you mouseover or click on a node of the corresponding interactions schematic (described below) depend upon whether the MMDB summary page is displaying the the static or interactive version of the molecular graphic .
INTERACTIVE DISPLAY : Once the interactive view loads, you can:
Click on the structure to stop the spin.
Click an icon in the corresponding interactions schematic to highlight molecule in both the schematic and the molecular graphic.
Right click on the structure to open a menu that allows you to control various aspects of the display (background color, display solvent accessible surface) and/or to "export image."
If you select " Export Image " from the menu, the molecular graphic will open in a separate window. In that separate window, you can right click on the exported image to use the browser's " save image as " function.
Reload the MMDB summary page to refresh the page and to reveal the "3D view" button again. Then repeat the steps above as many times as desired in order to save snapshots of the structure at the desired angles.
Each time you select "Export Image," a new, separate window will open, making it possible to view the structure from many angles simultaneously .
LAUNCH FULL iCn3D : Click the button to launch the advanced (full feature) version of iCn3D in another window.
The full feature version provides many additional controls for rendering, labeling, coloring, and saving the structure, as well as viewing corresponding sequence data.
Note that iCn3D will launch only if your web browser supports WebGL . If it doesn't, modify the settings in your web browser to enable WebGL, or, if needed, update your web browser to a newer version that supports WebGL. (See the WebGL site for more information about compatibility with various web browsers.)
The molecular components of the biological unit can include the following: Proteins , if present, are shown as circles: etc. Nucleotide sequences (DNA, RNA), if present, are shown as squares: etc. Chemicals , if present, are shown as diamonds: etc. Non-standard biopolymers , if present, are shown as parallelograms: (These are molecules such as nucleotide or protein sequences that contain a large percentage of non-standard residues.) etc. If any protein or nucleotide molecules in the structure were generated by applying transformations from crystallographic symmetry , their labels are shown as alphanumeric combinations (for example, or ), indicating the source molecule from which they were generated (to the left of the underscore bar ) and the copy number (to the right of the underscore bar). Chemicals that interact only with such molecules were also generated by applying transformations from crystallographic symmetry; their icon labels also include an underscore bar, with a number on either side of the underscore bar to indicate the source chemical and the copy number, respectively. The protein and nucleotide icons are scaled to show the relative sizes of those molecular components, so they are roughly comparable to each other based on molecular weight. All chemical icons are the same size.
Interactions among components are shown as lines , and an interaction is displayed only if there are at least 5 contacts at a distance of 4 Å or less between the heavy atoms of the molecules.
There is no meaning to the length of the lines in the interaction schematic. After the interactions are drawn, the diagram is flattened out to fit into the square, lengthening or shortening lines as needed.
Because of the latter thresholds, ions that are part of the biological unit may be missing from the interaction diagram, but they will be listed in the table of molecular components and interactions . Interactions for short peptides, or for molecule types other than protein, DNA/RNA, and chemical, are not calculated. Molecules, such as crystallization agents, etc., that are not part of the biologically active molecule are absent from both the interaction schematic and the molecular components list.
If the structure contains multiple biological units and you choose to display "all biological units," then the MMDB summary page for the structure will show a schematic cartoon (and corresponding molecular graphic ) for each one.
If the static molecular graphic is displayed:
Mouse over any node in the interactions schematic to view the molecule name.
Double click on a node in the interaction schematic to jump down to the corresponding part of the Molecules and Interactions table, which provides additional information about the molecule.
If the interactive molecular graphic is displayed:
Each node in the interaction schematic works as a toggle switch to highlight a molecule on/off.
Double click on a node in the schematic to highlight just that molecule in the 3D structure.
Double click on that molecule again in the interaction schematic to un-highlight the molecule and revert to the previous view of the 3D structure.
To highlight all molecules of the same type , click on the term " protein ," " nucleotide ," or " chemical " that appears at the bottom of the interactions schematic. (Click on the term again to toggle the highlight off, if desired.)
Download Structure Data (save 3D structure record)
Format : Cn3D , PDB , XML , JSON , PNG (image)
Data set : single 3D structure , all 3D structures , alpha carbons , PDB source file
Additional details : annotated illustration of download options , details about data saved in each file format : ASN.1 (Cn3D) , PDB , XML and JSON , save image of 3D structure , save structure components
Single 3D Structure
All 3D Structures
PDB source file
File Format : ASN.1 (Cn3D)
Data Set : your choice of Single 3D Structure , All 3D Structures , or Alpha Carbons .
Press the "Download" button.
Details about the data that are saved : (1) For X-ray crystallography or neutron diffraction of crystal structures: (a) If you have chosen to display the "first biological unit " or "all biological units" on the structure summary page, the "Download" operation will save the data for the specific biological unit displayed in the molecular graphic. The saved file will include sequence and spatial coordinate data that were present in the original PDB source file as well as data that were generated at NCBI by applying transformations from crystallographic symmetry , if applicable to that biological unit. (b) If you have selected the " asymmetric unit " display option, the "Download" operation will save the data that were present in the PDB source file, whether those data represented all, part, or multiple copies of a biological unit. The saved file will not include any data generated at NCBI by applying transformations from crystallographic symmetry. (2) For structures resolved by experimental methods other than X-ray crystallography or neutron diffraction of crystal structures, the "Download" operation will save the data that were provided by the author in the PDB source file. The concepts of asymmetric unit, biological units, and crystallographic symmetry do not apply to these structures. Note for both (1) and (2) above: The saved file may also include some modifications (relative to the original PDB source file) that occurred as a standard part of MMDB data processing. Some examples are provided below in the notes about PDB format.
File Format : PDB
Data Set : your choice of Single 3D Structure or Alpha Carbons
Press the "Download" button.
Details about the data that are saved : The PDB-formatted file that is downloaded when you select "Format: PDB" (in the " Download Structure Data " section of a structure summary page) has undergone content validation that is a standard part of data processing . Its content may therefore be somewhat different from that of the original PDB record. For example, some PDB records may have discontinous residue numbers, which exist in a free text field. MMDB assigns a consecutive series of positive integers to residues in biopolymers, using a numerical data field. In addition, MMDB resolves some discrepancies that might exist between the SEQRES records and the atomic coordinates. For example, if the structure's atomic coordinates reveal the presence of amino acids or nucleotides that are not listed in the SEQRES records of an original PDB file, MMDB will derive the biopolymer sequence from the atomic coordinates and not from the original SEQRES records. The derived biopolymer sequence will then appear in the MMDB record, and in the SEQRES records of the PDB-formatted file saved from the MMDB database. As a third example, the spans of secondary structures annotated on proteins might vary between PDB and MMDB records, as NCBI algorithmically identifies alpha helices and beta strands using purely geometric criteria and annotates the proteins using that information rather than the spans indicated in the original PDB file. Therefore, the content of a PDB-formatted record you save from an MMDB structure summary page may be different from the content of the original PDB file. To save an exact copy of the original PDB source file , click on the "Download" button that appears next to the PDB ID in the upper right hand corner of a structure summary page .
File Format : XML or JSON
Data Set : your choice of Single 3D Structure , All 3D Structures , or Alpha Carbons .
Press the "Download" button.
Details about the data that are saved : The "Download Structure Data" function on an MMDB summary page will act upon the biological unit(s) or asymmetric unit currently displayed in your browser window, and allow you to download the data file in varying levels of detail (also referred to as varying levels of complexity, or data sets ). It is also possible to save the data by using the Web API .
It is possible to save an image of the 3D structure in a number of ways, with varying degrees of customization possible:
To save the default image of a 3D structure :
Open the summary page page for the desired structure and view its molecular graphic
Use the Download Structure Data box that appears beside the molecular graphic , and select the options for Format : PNG (image)
Click on the " Download " button
To customize the viewing angle and/or background color, apply termini labels, or view solvent accessible surface, and then save the resulting snapshot :
Open the MMDB summary page page for the desired structure and view its molecular graphic , which by default shows a static image.
Click the "3D view" button in the lower left corner of the static image to load an interactive view of the structure. (The interactive view uses a simple version of iCn3D , NCBI's web-based 3D structure viewer, and requires a web browser that supports WebGL .)
Once the structure spins to the desired position, click on the structure to stop the spin
Right click to open a menu that allows you to control various aspects of the display and/or to "export image."
Select " Export Image " to open the view in a separate window.
Right click on the exported image to use the browser's " save image as " function.
Reload the page to reveal the "3D view" button again, then repeat the process as many times as desired in order to save snapshots of the structure at the desired angles.
Each time you select "Export Image," a separate window will open, making it possible to view the structure from many angles simultaneously .
To customize rendering style of the structure, highlight selected regions of the structure and/or corresponding sequence data , add labels , etc., and then save the state of the structure so you can reload it in the full-featured version of iCn3D in the future:
Open the MMDB summary page page for the desired structure and view its molecular graphic , which by default shows a static image.
Click the "3D view" button in the lower left corner of the static image to load an interactive view of the structure. (The interactive view uses a simple version of iCn3D , NCBI's web-based 3D structure viewer, and requires a web browser that supports WebGL .)
Then click the "full-featured 3D viewer" button to launch the full version of iCn3D in another window.
Use the various menu options to render the structure with the desired style, color, labels, viewing angle, etc.
Select the File/Save Files/State File to save the state of the structure in a file. (The file will be named "NXXX_statefile" by default, unless you select the "Save As" option in your browser, and file will be in *.txt format . The NXXX in the default filename represents the PDB ID of the structure.) You can then later open the statefile through the iCn3D "File/Open State" menu option.
Alternatively, you can select the File/Save Files/iCn3D PNG Image to save both the customized 3D image (as a *.png file ) and the state of the structure (as an *.html file ).
Specifically, the "File/Save Files/iCn3D PNG Image" option saves two files with a single action. The first file will be named NXXX_xxxxxxxxxxxxxxxxx.png, and the second will be named NXXX_xxxxxxxxxxxxxxxxx.html, where NXXX in the default filename represents the PDB ID of the structure and xxxxxxxxxxxxxxxxx is a hash tag that represents the customizations you made to the view. (Example filenames are: "1TUP-pgmMZ96uF2YhEsNc6.png" and "1TUP-pgmMZ96uF2YhEsNc6.html").
Additionally, the *.png file includes a " share URL " at the bottom. If a user opens that file in iCn3D, they will be able to see the structure in the same state in which they saved it, and it will be a live structure, so they can continue to view it interactively in iCn3D.
To render and save images using the wide range of controls that are available in NCBI's free standalone Cn3D structure-viewing program:
Open the MMDB summary page for the structure of interest.
In the " Download Structure Data " box, select " Format : ASN.1 (Cn3D) " and the desired " Data Set ," then press the " Download " button.
Open the file in Cn3D, where you can render, label, color the structure as desired. (To do this, install Cn3D on your computer, and if desired, configure it as a helper app.)
The Cn3D tutorial , provides detailed instructions on saving structures and images , including any special annotations you have made to the 3D structure view, such as adding labels or using specific drawing styles.
To render and save images using the controls provided by external 3D structure viewing programs that read PDB file format , such as Rasmol :
Open the MMDB summary page for the structure of interest.
In the " Download Structure Data " box, select " Format : PDB " and the desired " Data Set ," then press the " Download " button.
Open the file in any 3D structure viewer (e.g., Rasmol ) that reads PDB file format.
Render, label, color, and save the structure as desired, according to the instructions provided by the structure viewing program's help documentation
Molecular Components
Tabular list of molecular components
Column headers: label, count, molecule
Molecule label, count, & name
Protein annotation graphic
Domain families (protein classification)
Molecule label, count, & name
Thumbnail graphic
Molecule label, count, & name
Thumbnail graphic
Non-standard biopolymers
Molecule label, count, & name
Web API: URL format for displaying or saving a structure record: It is possible to view or save a 3D structure record by linking directly to it. The URL format, parameters, and allowable values, are as follows:
parameters and allowable values:
examples of URLs for displaying or saving 3D structure records :
Citing the Molecular Modeling Database:
Additional References: