PDB Statistics: Growth in Number of Unique Protein Sequences in Released PDB Structures (Cumulative) at Identity 95%

This chart shows the annual and cumulative numbers of protein sequences in released PDB structures. The chart can be viewed for a few different levels of sequence identity since the beginning of the PDB archive. The cumulative bars represent the growth in unique protein sequences (number of polymeric entities) across history. The yearly bars (dark blue) tell how many new protein sequences were added in a certain year.

Note: The total number of sequence clusters in the statistics table is linked to the sequence cluster group search result page. There is a default precision threshold in calculating the numbers for performance balance. So the statistics count may have a slight discrepancy compared to the actual non-redundant group search result when the result count approaches or goes above 10,000. The group search result page provides an accurate count. The statistics page provides the trend.

Chart is currently loading

Sequence cluster level:

YearNumber of New Protein SequencesTotal Number of Protein Sequences
19761313
19771023
1978326
1979632
1980436
19811046
19821864
19831175
19841186
19851298
19869107
198711118
198825143
198946189
199052241
199156297
199266363
1993232595
19944621,057
19953431,400
19964071,807
19975612,368
19987563,124
19998964,020
200010045,024
200110446,068
200211127,180
200315588,738
2004211910,857
2005234213,199
2006264415,843
2007296618,809
2008276021,569
2009281724,386
2010287027,256
2011263829,894
2012288432,778
2013309635,874
2014379539,669
2015314442,813
2016372446,537
2017390050,437
2018378354,220
2019409158,311
2020492863,239
2021451067,749
2022540673,155
2023515278,307
2024545783,764
2025627690,040
2026203492,074