# Introduction

**MiGA** stands for **Microbial Genome Atlas** and is designed to be the genome-equivalent of projects like the Ribosomal Database Project (RDP) and SILVA. In addition to serving as a repository, MiGA provides data and tools to classify and catalog microbial genomes and perform gene diversity analysis against publicly available reference genomes.

MiGA takes variety of inputs to help its users make taxonomic inferences about their draft or whole genome sequences. It uses a combination of the genome-aggregate Average Nucleotide Identity concept, or **ANI**, and the Average Amino-acid Identity, **AAI**, to taxonomically classify a genomic sequence against the genome sequences in its reference database. Part of MiGA’s strength lies in the >10,000 reference genomes that make up its database and an efficient heuristic algorithm to search a query genome against all of these genomes. The reference database is maintained in a way which allows it to be regularly updated and improved with minimal downtime. This ensures consistently improved accuracy of classification without a large impact to its users.

MiGA also follows best practices in genomic analyses that allow input sequences to represent isolate genomes, single-cell sequences, or metagenome-derived bins and contigs. These sequences can be searched against **NCBI RefSeq** and **Prokaryotic Genomes** for classification or be used as input for a clade project. The difference between and use cases for these different analyses are covered in subsequent sections of this document.

Users may also choose to explore an expanding base of environmental genomic data sets which can be used by beginners to familiarize themselves with the inner workings of MiGA before submitting their own sequences or by microbial ecologists interested in specific environments.

Although MiGA contains powerful taxonomic abilities, it is also meant to act as a robust database for a growing number of genomes of uncultivated microbes using binning or single-cell techniques. The overarching goal is for this database to act as a communal space for analysis and exploration of these known and unclassified genomes and allow users to answer questions such as:

1. Has my genome or single-cell been found previously in another sample or habitat?
2. What is the closest relative and likely taxonomic affiliation of my query genome(s) against (named) species represented by isolates or previously identified uncultivated taxa?
3. What sequence and gene-content distinguish my query genome from previously described taxa?
4. Can I perform a high-resolution phylogenetic and gene-content analysis of my collection of (unpublished) genome sequences, including draft genomes and SAGs?


# Create Account

New users of MiGA Online must register by going to [MiGA Online](http://microbial-genomes.org/) and setting up an account. Click on **Create User** and fill in the form with your name, email address and password. Then click on the **Create my account** button at the bottom of the page. Easy! and best of all, **free**!

After creating an account, when you log in to MiGA you will find a pull down menu in the upper right of the page with options for **Dashboard**, **Settings**, **Query datasets** and **Log out**. The Dashboard option will take you to your personal home page with links to some of the same things. The most important of these is **Query datasets** which will take you to a page with links to previous analyses you have made. The Settings link allows you to change your email address and password.

Click on MiGA in the upper left to go back to the MiGA home page to start a new analysis or explore genomes in the databases.


# Types of Analyses

In most cases, a MiGA analysis involves finding the closest match(es) between a genome and genomes in a reference database. The [MiGA home page](http://microbial-genomes.org/) presents four large square boxes titled **NCBI Prok**, **RefSeq**, **Preproc**, and **Clades**. In addition there is a large horizontal box at the bottom of the page entitled **Environments**. With the exception of Preproc, these selections represent the types of analyses you can make with MiGA. The difference between NCBI Prok, RefSeq, and Environments has to do with what database is searched against. Clades is a special case in which genomes are compared against a specific species only and offer higher intra-species resolution. Preproc is for processing input files like raw reads into contigs, or genome quality evaluation, without searching any database.

## Preproc

**Preproc** stands for preprocessing. If you click on the **Upload genome or metagenome** button in the Preproc box, you can submit raw or gzipped data in one of several formats and MiGA will attempt to assemble the reads into larger contigs and process these assemblies. This process uses a series of freely available tools (many from the [Enveomics Collection](http://enve-omics.ce.gatech.edu/enveomics)) to check quality, trim and match the reads.

Fill in the form, selecting the type of dataset and type of input from the pull-down menus. Select the files for upload by browsing to them on your computer and click the **Upload new dataset** button at the bottom of the page.

If the input was a metagenomic dataset, users may now choose to bin the contigs into population genomes (bins) using a program outside of MiGA, and then analyse the resulting bins using MiGA.

## RefSeq

The NCBI **RefSeq** database is a collection of taxonomically diverse, non-redundant and richly annotated sequences from high quality closed genomes (reference datasets), 1,942 as of March 2018. By searching genomes against this database, users can begin to understand their datasets. MiGA will identify closely related sequences which are annotated in the RefSeq database. Thus, RefSeq analysis is best suited to identify the closest related high-quality available genomes.

## NCBI Prok

The **NCBI Prok** database contains approximately 11,000 genomes from the Prokaryotic section of the NCBI Genome database. These genomes are from cultured microorganisms identified to species and are curated by the NCBI. This is the most comprehensive database available in MiGA, and is recommended for analysis of complete or draft genomes of isolates and single cell amplification, and when the goal is to classify the genome to the lowest rank possible.

You may also choose this database to analyze binned genomes, sometimes termed population genomes. These are assembled from metagenomes by binning, or grouping together longer contigs into sets which are hypothesized to be from the same organism. There are several programs for accomplishing this, and we currently do not have a recommendation for which program to use.

## Clades

Clade Projects involve the analysis of genome sequences assigned to the same or closely related species in order to assess gene content diversity and genetic relatedness among the genomes. That is, the genomes are compared among themselves to identify orthologous groups of proteins and provide descriptive statistics on the distribution of genes within the clade such as the size of the core gene set and the pangenome. Clade projects are ideal for micro-diversity and epidemiological studies, including plant and animal pathogens. MiGA Online offers an expanding collection of Clade Projects including all (complete and draft) genomes of *Bacillus cereus sensu lato*, *Marinobacter* spp, *Pelagibacter ubique*, and *Thaumarchaeota*. Developing new Clade Projects by external users is not currently possible but will be added to MiGA Online in the future. In the meantime, you may request the staff add your Clade Project to MiGA Online in the [Public Roadmap](http://roadmap.microbial-genomes.org/) or use the standalone version of MiGA available at <http://code.microbial-genomes.org/miga>.

## Environments

MiGA also allows searching against collections of genomes, including unclassified genomes, from specific environments or studies. The RefSoil collection is a subset of classified genomes from the RefSeq database that are associated with the soil environment. Other collections like the Tara Ocean project MAGS are made up of unclassified genomes. The goal of searching a query genome against these databases is not to precisely classify it, but rather to determine if the query genome has been seen before in these specialized genome collections, and how similar it is to its closest match in these collections. Besides RefSoil and Tara Oceans MAGs collections accessible from MiGA's home page, additional environmental and study collections are available by clicking on the **Projects** link in the top navigation bar.


# Understanding Results

Access results of a query by selecting **Dashboard** from the drop down menu in the upper right of any page, click on **Query data set**, and then click on the name of your query. You will be presented with a page similar to that below, perhaps with **Distances** already open. The headings shown depend on the analysis made. **MyTaxa Scan** and **Quality** provide information related to the quality of genomes (query and reference) and will be discussed in the Quality section.

![directory structure](/files/-LBP6yax72gFkw_AmDpR)

You can open and close each of the result sections (Distances, Ribosomal RNA, Quality, Gene prediction, and Assembly) by clicking on the section title.

## Distance

If you click on the Distance bar, you get something that looks like this:

![directory structure](/files/-LBP6ybfVJy1wMNOZqVi)

The closest matches in the data base are displayed under Distance, the classification with probabilities under Taxonomic classification, and the probability that the query genome belongs to a new taxon under Taxonomic novelty. Links under Genomic relatedness give the relationship to closest matches in the form of a tree and a table of AAI percentages.

Relationships are calculated on the basis of ANI (average nucleotide identity) and AAI (average amino acid identity). MiGA makes inferences based on the obtained values and predetermined thresholds for each taxonomic rank. For instance, if a given sequence shares at least 95% ANI with a reference genome, there is a high likelihood that it should be assigned to the same species. For more divergent query sequences AAI, is used. Genome sequences are hypothesized to be from the same genus if they have an AAI above 65%. If the sequences show AAI between 45 and 65%, they are hypothesized to be from the same family. Further information on this approach and thresholds used can be found by clicking on the information icons (blue circles with i's in them) next to each title.

MiGA’s P-values for each assignment to a taxonomic rank reflect the confidence for the assignment given the distribution of AAI or ANI values among all genomes grouped at the same rank, e.g., all members of a species for species-level assignments and where within this distribution the corresponding ANI values between the query sequence and its best matching reference genome fall. That is why MiGA will not make any inferences on p-values over 0.5 as that indicates a greater possibility that the identity value obtained for the query genome lies outside of the distribution of values for the taxonomic rank in evaluation than inside of it.

MiGA also provides a confidence (p-value) that the genome is novel at a specific rank that is essentially the inverse of the p-values of classification in the previous step. Here you see the classifications with the higher p-value from before expressing the higher confidence that they do belong to novel classifications. For our sample genome, there is high confidence that the sequence belongs to a species, genus, and family not represented by the genomes in the reference database.

Below this section is an AAI table which shows AAI values for the query sequence against the closest related genomes in the database.

![directory structure](/files/-LBP6yciie0Pq3Qau4pp)

In this example, we can see that with an AAI of 73.13% a confident assignment to genus is made. The other reference genomes listed have AAI' of \~55% and belong to a different family (*Vibrionaceae*). Closer examination of the ANI table allows the user to infer species level similarity.

## Ribosomal RNA

Analysis of the 16S ribosomal RNA gene is also offered when the gene sequence is detected. Classify it with the RDP classifier by clicking on the link or download the sequences by clicking on the download icon (downward pointing arrow).

![directory structure](/files/-LBP6ygTSbKpED39Hv50)

The thresholds on sequence similarity are based on the thresholds established by the Ribosomal Database Project, or RDP:

| Same    | % Similarity | Novel  | % Similarity |
| ------- | ------------ | ------ | ------------ |
| Family  | 92-95        | Domain | < 75         |
| Genus   | 95-98.6      | Phylum | 75-83        |
| Species | 95-100       | Order  | 83-86        |
| -       | -            | Class  | 86-89        |
| -       | -            | Family | 89-92        |

To summarize, MiGA offers a variety of graphical outputs in order to help users understand their sequences and classification results. Understanding the goals and limitations of each analysis is essential to accurate interpretation. Often, confident classification can only be made at specific ranks of taxonomy and not others. MiGA’s classification strength will continue to grow as more reference genomes become available. So please deposit your new genome sequences to MiGA!


# Assessing Quality

## Summary Statistic

Clicking on the Quality heading on the Results page accesses a summary statistic related to the quality of the query genomic sequence or a genome in the reference database. This summary statistic relates completeness and contamination. Completeness is measured by the presence of 111 single-copy genes which are observed across almost all prokaryotic genomes, while contamination is measured by the frequency at which these genes are present in more than one copy. The display below includes the "Full report" obtained by clicking on the button so named.

![directory structure](/files/-LBP6zAqC89rRZsbA1o8)

Our example genome is a high-quality assembly as it contains 106 out of 111 single copy genes. It also shows redundancy of three of these genes, with each one showing two copies. Quality scores are then calculated as completeness percentage minus five times contamination percentage. This frequency of duplicated genes can also be referred to as chimerism as it usually but not necessarily always is indicative of a chimeric combination between the sequence of the main genome and other highly similar but distinct sequences from other organisms. Chimerism can arise during binning or can reflect actual DNA contamination, depending on the source. It is important to note that this is a universal attempt at quality assessment. Further analysis via techniques such as single cell amplified genomes can confirm if the multiple copies of traditional single-copy genes are real or the product of chimerism. This chimerism can also be explored further via the MyTaxa scan tool.

## MyTaxa Scan

MyTaxa scan provides another means of quality checking. This is a CPU intensive task and an optional analysis that a user has to request by checking the corresponding box when submitting the analysis. MyTaxa scan scans the whole input genome in windows of ten genes and determines the phylogenetic origin of each window. MyTaxa then flags windows that deviate significantly from the genome average. When the analysis is finished, a PDF report is generated which provides information on the distance of each window from the average of the genome and options to check the origin of each window, download the corresponding genes, or other options. Windows that deviate from the genome average may represent cases of recent horizontal gene transfer or contamination. For further details on MyTaxa scan, refer to the original MyTaxa publication.&#x20;

The My taxa scan pdf for the high quality genome above shows little evidence of contamination:

![directory structure](/files/-LBP6zBWBtMtcgm-WTdF)

My taxa results for a low quality genome would look more like this:

![directory structure](/files/-LBP6zC5IRxoeV9o9uGa)

## Gene Prediction & Assembly

Finally, doenload links to information on gene prediction and assembly length of the input genomic sequence is also included on the results page.

## Summary

By utilizing the available information generated by MiGA users can thoroughly evaluate the quality of their assembled genomic sequences and even make further inferences about the success of different assembly or binning pipelines. In the end, areas of chimerism should always be manually examined and may not always indicate diminished genome quality; they may be cases of recent horizontal gene transfer or other naturally occurring genetic deviation.


