DOE Joint Genome Institute

  • COVID-19
  • About Us
  • Contact Us
  • Our Science
    • DOE Mission Areas
    • Science Programs
    • Science Highlights
    • Scientists
    A photo of Great Boiling Spring in the forefront with mountains in the background.
    New Research Finds Flagella in the Terrestrial Roots of Marine Bacteria
    Scientists have discovered flagella in an unexpected place: hot spring-dwelling bacteria. Research shows that flagella were lost in other forms of Chloroflexota that adapted to marine environments hundreds of millions of years ago

    More

    Sphagnum fallax (Image courtesy of Jonathan Shaw, Duke University)
    Sequencing Sphagnum Leads to Discovery of Sex Chromosomes
    Boggy peatlands are primarily made up of sphagnum mosses. New research illuminates the significant role sex plays in how the moss grows, stores carbon and responds to stress.

    More

    A vertical tree stump outdoors with about a dozen shiitake mushrooms sprouting from its surface.
    Tracing the Evolution of Shiitake Mushrooms
    Understanding Lentinula genomes and their evolution could provide strategies for converting plant waste into sugars for biofuel production. Additionally, these fungi play a role in the global carbon cycle.

    More

  • Our Projects
    • Search JGI Projects
    • DOE Metrics/Statistics
    • Approved User Proposals
    • Legacy Projects
    A panoramic view of a lake reflecting a granite mountain.
    Genome Insider: Methane Makers in Yosemite’s Lakes
    Meet researchers who sampled the microbial communities living in the mountaintop lakes of the Sierra Nevada mountains to see how climate change affects freshwater ecosystems, and how those ecosystems work.

    Listen

    A light green shrub with spiny leaves, up close.
    Genome Insider: A Shrubbier Version of Rubber
    Hear from the consortium working on understanding the guayule plant's genome, which could lead to an improved natural rubber plant.

    Listen

    An underwater photo with a blue backdrop and a mineral formation erupting black, mineral-rich fluid.
    Shedding Light on Diversity in the Deep Sea
    Hydrothermal vents are too deep in the sea to get any sunlight, and yet they still support a unique array of microorganisms adapted to living in extreme conditions.Researchers studied microbial communities at these vents across five oceanic regions.

    More

  • Data & Tools
    • IMG
    • Data Portal
    • MycoCosm
    • PhycoCosm
    • Phytozome
    • GOLD
    Illustration of a magnifying glass identifying viruses and plasmids.
    geNomad: Rapidly identifying mobile genetic elements
    geNomad builds on two standard techniques for identifying viruses and plasmids. It can identify millions of new viruses and plasmids quickly, even in massive datasets.

    More

    LISTEN: Natural Prodcast on the SMC
    Check out the JGI Secondary Metabolism Collaboratory (SMC), a new data portal for natural product biosynthetic gene clusters and meet Prodcast cohost Jackie Winter.

    More

    iPHoP image (Simon Roux)
    iPHoP: A Matchmaker for Phages and their Hosts
    Building on existing virus-host prediction approaches, a new tool combines and evaluates multiple predictions to reliably match viruses with their archaea and bacteria hosts.

    More

  • User Programs
    • Calls for Proposals
    • Special Initiatives & Programs
    • Product Offerings
    • User Support
    • Policies
    • Submit a Proposal
    A series of headshots: From left to right: [above] Olivia Ahern, Adriana Corales, Hugh Cross, Megan DeMarche, Joanne Emerson, Matthew Hudson, Megan Keller and Julia Kelliher; [below] Vassili Kouvelis, Seppe Kuehn, Tesfaye Mengiste, Egbert Schwartz, Hannah Schulman, Bram Stone and Jana Voriskova
    2024 awardees for JGI Community Science Program annual call
    Learn about the 15 projects selected through the JGI's FY2024 Community Science Program Annual Call - and the investigators behind them.

    More

    From left to right: [above] Emma Bell, Mallory Choudoir, Sneha Couvillion, Tobin Hammer, Christina Hazard, Rachel Mackelprang, Brook Moyers, Mei, Ran,; [below] Benjamin Peterson, Dacheng Ren, Allison Rober, Neal Scott, Chikae Tatsumi, Vojtech Tlaskal, Fernando Torralbo, Luis Felipe Valdez-Nuñez
    JGI announces second round of 2023 New Investigator awardees
    Learn about accepted proposals on novel research projects — aligned with DOE missions and from PIs who have not led any previously-accepted proposals.

    More

    Green plant matter grows from the top, with the area just beneath the surface also visible as soil, root systems and a fuzzy white substance surrounding them.
    Supercharging SIP in the Fungal Hyphosphere
    Applying high-throughput stable isotope probing to the study of a particular fungi, researchers identified novel interactions between bacteria and the fungi.

    More

  • News & Publications
    • News
    • Blog
    • Podcasts
    • Webinars
    • Publications
    • Newsletter
    • Logos and Templates
    • Photos
    A Decade On: The JGI-UC Merced Genomics Internship Program
    What started with 2 interns in the summer of 2014 has grown to 75 student alumni. Mentors, students and others gathered to reflect on the benefits of this partnership.

    More

    A tiled collage of square photos of different plants - soybeans, and sorghum, for example.
    A Collaboration to Improve Plant Genome Annotations Across Species
    The JGI Plant Gene Atlas is an updateable transcriptome resource spanning diverse plant species. The project spans 15 years and involves more than 17 research groups.

    More

    2022 JGI-UC Merced interns (Thor Swift/Berkeley Lab)
    Exploring Possibilities: 2022 JGI-UC Merced Interns
    The 2022 UC Merced intern cohort share how their summer internship experiences have influenced their careers in science.

    More

News & Publications
Home › News Releases › Defining Quality Virus Data(sets)

December 17, 2018

Defining Quality Virus Data(sets)

International consortium offers guidelines, best practices for characterizing uncultivated viruses.

Artist rendering of genome standards being applied to deciphering the extensive diversity of viruses. (Illustration by Leah Pantéa)

Artist rendering of genome standards being applied to deciphering the extensive diversity of viruses. (Illustration by Leah Pantéa)

Microbes in, on and around the planet are said to outnumber the stars in the Milky Way Galaxy. The total number of viruses is expected to vastly exceed even that calculation.

While many viruses remain unknown and uncultivated, advances in genome sequencing and analyses have allowed researchers to identify more than 750,000 uncultivated virus genomes from metagenomic and metatranscriptomic data sets. In IMG/VR, a database for virus sequences established and maintained by researchers at the U.S. Department of Energy (DOE) Joint Genome Institute (JGI), a DOE Office of Science User Facility, the viral diversity available has tripled within a single year.

As more and more researchers continue to assemble new genome sequences of uncultivated viruses, JGI researchers led a community effort to develop guidelines and best practices for defining virus data quality. In a report published December 17, 2018, in Nature Biotechnology, JGI partnered with a number of virus experts; as well as representatives from the Genomic Standards Consortium (GSC), an open-membership working body that engages the research community in the standards development process; and the International Committee on Taxonomy of Viruses, the premier authority on the official taxonomy of viruses which is currently re-evaluating virus classification based on sequence-based information.

Guidelines for Quality and Analyses

“Viruses are critical components of every microbial ecosystem. The JGI is especially interested in developing standards for virus genomes because we generate much of this data ourselves,” said JGI research scientist and first author Simon Roux. “We are part of a small group of researchers who have scrutinized these data at length, have seen the metrics, and can provide guidance to help determine data quality. Additionally, in this paper, we’ve tried to provide not just standards, but also outline what type of analyses can be performed on these data, to help researchers who want to characterize their own novel viruses.”

Cultured viruses already have their own data quality standards, but these cannot be directly applied to uncultured viruses, whose sequences are often incomplete and for which some properties can only be predicted indirectly using computational approaches.

“The uncultivated virus genome community has come together to define what is important to report and valuable to the research community,” said GSC President Lynn Schriml of the Institute of Genome Sciences at University of Maryland School of Medicine. The GSC includes representatives from the National Center for Biotechnology Information (NCBI), the European Bioinformatics Institute, and the DNA Data Bank of Japan (DDBJ), who also collaborated on this article.

Categories of Virus Genome Quality

In the paper, Roux and his colleagues outlined the minimum amount of information for an uncultivated virus genome, including the source, methods of identification of the virus genome, and data quality. The JGI has previously developed standards for the minimum metadata to be reported with single amplified genomes (SAGs) and metagenome-assembled genomes (MAGs) submitted to public databases.

“The tremendous growth of virus sequence data, and microbiome data in general, necessitates robust standards and data quality metrics to allow the research community to leverage this data for comparative analyses,” said JGI Metagenome Program head and study senior author Emiley Eloe-Fadrosh. “By establishing and promoting ‘best practices,’ the research community has the opportunity to break down barriers of data accessibility and reusability, thereby amplifying the research beyond the initial project scope.”

The team proposed three categories of genome quality. “Genome fragments” are comprised of single or multiple fragments that are predicted to be less than 90 percent complete, or have no estimated genome size, and are minimally annotated. A “high-quality draft genome” is estimated to represent 90 percent or more of the complete expected genome sequence, in fragment(s) where any gaps span mostly repetitive regions. Finally, a “finished genome” would include both a complete genome comprised of a single contiguous sequence without gaps, and extensive annotation.

“If you’re going to build a standard,” Schriml noted, “it is essential to discuss what should be represented with the research community, taxonomists and database providers and to integrate these data needs into the standard.”  Schriml added journals have also started endorsing the application of the GSC’s “Minimum Information about any (X) Sequence (MIxS)” guidelines, the umbrella under which the uncultivated virus genome standards and other similar community efforts reside. The GSC tracks the adoption of these standards developed over the past decade using records uploaded to the BioSample database. These records reflect individual samples collected, sequenced and annotated, and Schriml said that nearly 450,000 BioSample records currently reference MIxS guidelines, up from 326,000 records tracked in the spring.

 

Reference: Roux S et al. Minimum Information about an Uncultivated Virus Genome (MIUViG). Nature Biotechnology. 17 December 2018. https://doi.org/10.1038/nbt.4306

 

Byline: Massie Santos Ballon

Share this:

  • Click to share on Facebook (Opens in new window)
  • Click to share on LinkedIn (Opens in new window)
  • Click to share on Pinterest (Opens in new window)
  • Click to share on Twitter (Opens in new window)
  • Click to print (Opens in new window)

The U.S. Department of Energy Joint Genome Institute, a DOE Office of Science User Facility at Lawrence Berkeley National Laboratory, is committed to advancing genomics in support of DOE missions related to clean energy generation and environmental characterization and cleanup. JGI provides integrated high-throughput sequencing and computational analysis that enable systems-based scientific approaches to these challenges. Follow @jgi on Twitter.

DOE’s Office of Science is the largest supporter of basic research in the physical sciences in the United States, and is working to address some of the most pressing challenges of our time. For more information, please visit science.energy.gov.

Filed Under: News Releases

More topics:

  • COVID-19 Status
  • News
  • Science Highlights
  • Blog
  • Webinars
  • CSP Plans
  • Featured Profiles

Related Content:

JGI announces 2023 CSP Functional Genomics awardees

Digital index card with JGI logo reads: Community Science Program (FY23) Congratulations to our CSP Functional Genomics recipients! Picture from left to right: (top) Thom Booth, Gabriel Castrillo, Han Li; (bottom) Jorge A. Marchand, Emre Özdemir, Fong Tian Wong

Researching and Solving Real-World Problems with the 2023 JGI-UC Merced Interns

2023 JGI-UC Merced interns (Zhong Wang/Berkeley Lab)

RECAP: Multi-Omic Journeys with 2023 JGI Annual Meeting Keynotes

Bruce Hungate stands at a podium and gesticulates as he discusses microbes.

For the Tiniest Archaea, A Genomic Switch of Friend or Foe

A grey microscopy photo taken at micron-scale. Microbes shown are small, round and slightly spiky in shape.

Doubling Down on Known Protein Families

An illustration of a microscope emitting a beam of light that hits a small, nondescript item.

The JGI announces 2024 awardees for our Community Science Program annual call

A series of headshots: From left to right: [above] Olivia Ahern, Adriana Corales, Hugh Cross, Megan DeMarche, Joanne Emerson, Matthew Hudson, Megan Keller and Julia Kelliher; [below] Vassili Kouvelis, Seppe Kuehn, Tesfaye Mengiste, Egbert Schwartz, Hannah Schulman, Bram Stone and Jana Voriskova
  • Careers
  • Contact Us
  • Events
  • User Meeting
  • MGM Workshops
  • Internal
  • Disclaimer
  • Credits
  • Policies
  • Emergency Info
  • Accessibility / Section 508 Statement
  • Flickr
  • LinkedIn
  • RSS
  • Twitter
  • YouTube
Lawrence Berkeley National Lab Biosciences Area
A project of the US Department of Energy, Office of Science

JGI is a DOE Office of Science User Facility managed by Lawrence Berkeley National Laboratory

© 1997-2023 The Regents of the University of California