Practical 1a. Phylogeny-Based Analysis of DNA Barcode Sequences
This tutorial provides step-by-step instructions for phylogenetic placement and analysis using T-BAS and DeCIFR.
Part 1 — Taxonomy Resolution with Multiple Databases
Integrate phylogenetic placement with multiple reference databases to improve taxonomic resolution and evaluate assignment support.
Why Phylogenetic Placement?
Why it matters
- Places unknown ITS sequences into an evolutionary context.
- More informative than similarity searches alone.
- Estimates placement confidence.
- Supports downstream diversity analyses.
Key outputs
- Evolutionary placement
- Likelihood weights
- Interactive tree
Goal: Determine where unknown sequences belong on the fungal tree of life.
Screenshot guide

Step 1. Select a Reference Tree.
What you will do
- Select the Fungi v3 reference tree.
- Download the example datasets.
- Prepare the ITS ASV FASTA file.
Observe
- Choose the reference matching the ITS marker.
- The tutorial uses Fungi v3.
- Next: Upload the ITS sequences.
Screenshot guide

Practical 1a. Phylogenetic Placement of DNA Barcode Sequences.
- Click here to select reference tree
Screenshot guide

Select Fungi v3 for phylogenetic placement of ITS reads.
- Target reference tree for placement
Screenshot guide

Download example files.
- Click on examples
Screenshot guide

Download the metadata and reference paper used in this tutorial. A ZIP archive containing all files is also available.
Screenshot guide

Step 2. Upload ITS Sequences.
What you will do
- Upload the ASV FASTA file.
- Enable ITS filtering.
- Select RDP, FUNGuild and NCBI RefSeq.
Common mistakes
- Wrong FASTA format
- Mixed non-ITS sequences
- Wrong reference tree
- ITSx identifies non-ITS sequences before placement.
Screenshot guide

Drag the ASV fasta file into the unknown query box.
Screenshot guide

Select filter unknowns and generate UNITE report.
Screenshot guide

Select RDP Classifier, FUNGuild, and NCBI ITS RefSeq databases.
Screenshot guide

Step 3. Configure EPA-ng.
What you will do
- Choose EPA-ng.
- Provide a run name.
- Select the ITS locus.
- Submit the analysis.
Runtime
- ~1 hour
- Monitor progress
- Email notification
- EPA-ng is the longest-running step in the tutorial.
Screenshot guide

Use EPA-ng (Evolutionary Placement Algorithm – Next Generation).
Screenshot guide

Provide a label for the run for future reference.
Screenshot guide

Select the ITS locus in the pull-down menu.
Screenshot guide

Step 4. Review the ITSx Report.
What you will do
- Review the ITSx quality-control report.
- Identify ITS-only sequences.
- Trim or remove sequences, if necessary, before placement.
Observe
- Most sequences should contain the expected ITS region.
- Problematic sequences can be excluded before analysis.
- Quality control improves the accuracy of downstream phylogenetic placement.
Screenshot guide

Verify ITS-Only Sequences using ITSx.
Screenshot guide

ITSx report reveals ITS1-only sequences.
Screenshot guide

Have the option to submit, cancel or trim sequences.
- Because these are ITS1-only sequences, click Submit to continue
Screenshot guide

Step 5. Monitor the EPA-ng Analysis.
What you will do
- Monitor the progress bar.
- Wait for the analysis to complete (~1 hour).
- Open the completed run directory.
- View the summary report.
Expected outputs
- Run summary
- Placement statistics
- Interactive tree
- EPA-ng sends an email when the analysis is complete.
Screenshot guide

Progress bar displays status of run.
- Takes about an hour to finish
Screenshot guide

Summary reports at end of run and option to view tree.
- Click to view tree
Screenshot guide

Completed run is also announced via email.
Screenshot guide

Step 6. Explore the Interactive Tree.
What you will do
- Open the interactive phylogenetic tree.
- Display taxonomic groups.
- Locate the placed query sequences.
- Inspect nearby reference taxa.
Observe
- Unknown sequences are placed within an evolutionary context.
- Closely related taxa provide biological interpretation.
- Phylogenetic placement is more informative than a simple database match.
Screenshot guide

Tree view in T-BAS showing Class and Unknowns shaded in grey.
Screenshot guide

Step 7. Evaluate Placement Confidence.
What you will do
- Color branches using EPA-ng LWR support.
- Compare EPA-ng placement support with gappa taxonomic assignment support.
- Inspect individual placements.
Interpretation
- High LWR = confident placement.
- Low LWR may indicate ambiguous taxonomy or missing references.
- Confidence values are essential for interpreting placement results.
Screenshot guide

Display EPA-ng and gappa LWR Support.
- Select EPA and gappa
Screenshot guide

Interpret EPA-ng and gappa LWR Support.
- High placement support does not necessarily mean high-confidence species identification.
EPA-ng LWR
- Placement support
- Where does the query place?
gappa LWR
- Taxonomic assignment support
- Do the plausible placements support the same taxonomic assignment?
Interpret the two values together
- High EPA + high gappa: Confident placement and confident taxonomic assignment.
- High EPA + low gappa: Placement may be strong, but species-level assignment remains uncertain.
Take-home message: Strong phylogenetic placement ≠ automatically strong species assignment.
Screenshot guide

Step 8. Inspect Sequence Alignments.
What you will do
- Select a focal clade.
- Open the multiple sequence alignment viewer.
- Compare query sequences with nearby references.
Observe
- Sequence similarity
- Conserved positions
- Potential alignment problems
- Alignment inspection helps validate unusual placements.
Screenshot guide

Select the Focal Clade.
- Uncheck Use branch length.
- Set Circle diameter multiplier = 20.
- Locate the focal clade.
- Click the node defining the clade to select it.
- Next: zoom in to make the focal clade easier to select.
Screenshot guide

Zoom In to Select the Focal Clade.
- Use + to zoom into the focal region.
- Locate the branch defining the clade of interest.
- Click the node defining the clade to select it.
Screenshot guide

View the Alignment for the Selected Clade.
- Click View to open the alignment
Screenshot guide

Compare query sequences with nearby references.
- Look for:
- Sequence similarity
- Conserved and variable sites
- Unexpected gaps
Screenshot guide

In-Class Discussion 1.
- Interpreting Phylogenetic Placement
Before coming to class, think about
- Where did most of your ITS sequences place on the reference tree?
- Which placements had high confidence? Which were uncertain?
- Why might some sequences have ambiguous placements?
- What advantages does phylogenetic placement provide over a BLAST similarity search?
Learning goal
- Explain how phylogenetic placement provides an evolutionary framework for identifying unknown DNA barcode sequences.
- Complete Practical 1A before class and bring your results for discussion.
Screenshot guide

Part 2 — Microbial Community Diversity
Analyze phylogenetic and community diversity using complementary alpha- and beta-diversity approaches.
Why Analyze Microbial Community Diversity?
Why analyze diversity?
- Compare microbial communities using complementary phylogenetic and abundance-based approaches.
- Measure diversity within and differences among samples.
- Quantify changes across treatments and sampling dates.
- Link phylogeny with ecological interpretation.
Two complementary views of diversity
Alpha diversity — within a sample
- Faith's PD — phylogeny-based diversity
- How much evolutionary diversity is present?
Beta diversity — among samples
- UniFrac — phylogeny-based dissimilarity
- Bray–Curtis — abundance-based dissimilarity
- NMDS (Non-metric Multidimensional Scaling) — visualizes Bray–Curtis dissimilarities
- How different are microbial communities?
Take-home: Alpha diversity describes diversity within samples; beta diversity describes differences in community composition among samples.
Screenshot guide

Step 9. Launch UniFrac Analysis.
What you will do
- Open the UniFrac analysis.
- Upload the ASV count table.
- Upload the sample metadata.
- Submit the analysis.
Runtime
- ~1 minute
- Results stored in the run directory.
- UniFrac compares communities using branch lengths on the reference tree.
Screenshot guide

Perform diversity analysis using UniFrac.
- Click on UniFrac
Screenshot guide

Upload the ASV counts file.
Screenshot guide

Upload the sample metadata file with experimental attributes.
- Include a note about the run and click submit
Screenshot guide

Wait for the run to complete.
- Takes a minute to finish
Screenshot guide

Step 10. Review UniFrac Results.
What you will do
- Open the completed run directory.
- Locate the Faith_pd and matplotlib folders.
- Review tables and publication-ready figures.
Key outputs
- Faith_pd/: statistics
- matplotlib/: figures
- Most analysis outputs are organized automatically.
Screenshot guide

Click on the run directory to view the results.
Screenshot guide

Step 11. Interpret Faith's PD.
What you will do
- Open the Faith’s PD figure.
- Compare phylogenetic diversity across sampling dates.
- Examine the overall temporal pattern.
Observe
- Faith’s PD generally increases across the 2022 sampling dates.
- Later samples tend to have greater phylogenetic diversity.
- Faith's PD measures the total evolutionary diversity within each sample.
Screenshot guide

Locate the Faith_pd/ Output Folder.
- Faith's PD results and figures are saved in Faith_pd/. Open this folder to examine phylogenetic alpha diversity.
- Click Faith_pd/Faith's PD resultsand figures
Screenshot guide

Visualize Faith's Phylogenetic Diversity.
Screenshot guide

In-Class Discussion 2.
- Interpreting Phylogenetic Diversity
Before coming to class, think about
- Which sampling dates have the greatest phylogenetic diversity?
- Do the results agree with your expectations?
- What biological processes could explain the observed pattern?
- Why might Faith's PD differ from species richness?
- Complete Practical 1A before class and bring your results for discussion.
Learning goal
- Interpret changes in phylogenetic diversity and relate them to microbial community dynamics.
Screenshot guide

Step 12. Interpret NMDS.
What you will do
- Open the Bray–Curtis NMDS plot.
- Compare clustering by survey date.
- Relate ordination to Faith's PD.
Observe
- Samples separate over time.
- Community composition shifts through the season.
- Ordination complements diversity metrics by revealing community structure.
Screenshot guide

Locate the matplotlib/ Output Folder.
- Diversity results are organized into several output folders. Open matplotlib/ to view publication-ready figures with statistical annotations.
- Click matplotlib/Publication-ready figures with statistical annotations.
Screenshot guide

Visualize Temporal Changes in Microbial Community Composition.
Screenshot guide

In-Class Discussion 3.
- Comparing Microbial Communities
Before coming to class, think about
- Which samples cluster together?
- Which sampling dates are most distinct?
- Do the NMDS and Faith's PD results tell the same story?
- What ecological processes might explain these community shifts?
- Complete Practical 1A before class and bring your results for discussion.
Learning goal
- Interpret differences in community composition using beta-diversity analyses.
Screenshot guide

Part 3 — Taxonomy Resolution with Multiple Databases
Integrate phylogenetic placement with multiple reference databases to improve taxonomic resolution and evaluate assignment support.
Why Taxonomy Resolution Is Needed.
- Phylogenetic placement accurately places sequences on a reference tree, but resolution is limited by the taxa represented in that tree.
- Reference databases (UNITE, RDP, and NCBI RefSeq) provide broader coverage but can contain conflicting or incomplete classifications.
- The Taxonomy Resolver integrates T-BAS placements with multiple databases to recover the best-supported taxonomy.
Take-home message: Integrating phylogenetic placement with multiple reference databases produces more robust taxonomic assignments than either approach alone.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Launch the Taxonomy Resolver
- Integrate T-BAS placements with UNITE, RDP, and NCBI RefSeq to recover the best-supported taxonomy beyond the limits of any single reference tree.
Screenshot guide

Step 13. Configure the Taxonomy Resolver.
What you will do
- Enter the T-BAS accession.
- Upload the ASV count table.
- Upload the metadata table.
- Choose the recommended taxonomy strategy.
Required inputs
- Placement accession
- ASV counts
- Metadata
- Read mode
- These inputs provide the information required for downstream analyses.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Enter T-BAS accession
- Upload ASV count table
- Specify read mode
- Upload metadata table
- Choose the recommended taxonomy strategy
Screenshot guide

Step 14. Enable Downstream Analyses.
What you will do
- Enable DESeq2.
- Enable Faith's PD.
- Complete all required fields.
- Submit the analysis.
Runtime
- ~1 minute
- Results saved under analysis_output/
- The Taxonomy Resolver can automatically launch downstream analyses.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Check “Run” to enable DESeq2 analysis
- Click the warning banner to step through the required fields.
- Check “Run” to enable Faith's PD analysis
- Select the downstream analyses to perform after taxonomy resolution.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Complete the required fields
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Complete the three required fields (red *). For repeated-measures studies, also specify the block column (blue *). Click the warning banner to reveal any remaining required fields.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Complete the highlighted required fields
- Continue clicking the warning banner until all required fields have been completed.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Confirm the required grouping fields:
- Treatment and Survey_Date
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- Enter a run name, then click Extract taxonomy
Screenshot guide

Step 15. Explore the Results.
What you will do
- Open the consensus taxonomy table.
- Review the best-supported taxonomy.
- Locate DESeq2 and Faith's PD output folders.
- Inspect publication-ready figures.
Key outputs
- Consensus taxonomy
- DESeq2 figures
- Faith's PD figures
- Downstream Taxonomy Resolver outputs are written to analysis_output/; placement/support reports remain in the run directory.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
- The analysis typically completes in ~1 minute
- Open the run directory to view the results
Screenshot guide

Example Output Paths.
The placement/support table is in the run directory; downstream Taxonomy Resolver results are under analysis_output/.
Screenshot guide

Step 16. Interpret the Results.
- Locate the placement-support report generated by the Taxonomy Resolver.
- Interpret EPA-ng and gappa likelihood weight ratio (LWR) values.
- Review the consensus taxonomic assignment.
- Interpret downstream diversity and differential-abundance results.
Questions to ask
- Is the phylogenetic placement well supported?
- Is the taxonomic assignment well supported?
- Do placement and taxonomic support agree?
Screenshot guide

Examine Placement and Taxonomic Support.
Output file: assignments_report_pretty_withgappa(T-BAS run ID).csv
In this example: assignments_report_pretty_withgappa6FAEQW0E.csv
- Examine the placement report to evaluate placement and taxonomic support for each ASV.
Screenshot guide

Interpret the Four Placement-Support Columns.
EPA Likelihood Weight Ratio (LWR)
Relative support among alternative phylogenetic placements of the query.
gappa Likelihood Weight Ratio (LWR)
Likelihood weight associated with the taxonomic assignment.
gappa fraction of placements (fract)
Fraction of candidate placements supporting that taxonomic assignment.
gappa accumulated LWR (aLWR)
Accumulated likelihood weight supporting that taxonomic assignment.
Key distinction: EPA-ng evaluates where the sequence can be placed on the reference tree; gappa summarizes how those placements support taxonomy.
Screenshot guide

How to Read EPA-ng and gappa Support.
- EPA-ng LWR near 1.0: placement support is concentrated on one branch.
- Split EPA-ng LWR values: multiple placements remain plausible.
- High gappa support: the placements consistently support the same taxonomic assignment.
- Low gappa support: taxonomic assignment remains uncertain even when a placement can be made.
- Interpret both together:High EPA-ng + high gappa support → confident placement and taxonomySplit/low support → treat the assignment cautiously
Screenshot guide

From Placement Support to Consensus Taxonomy.
- Use EPA-ng LWR to assess confidence in phylogenetic placement.
- Use gappa support to assess confidence in the associated taxonomy.
- Use the Taxonomy Resolver consensus table for the final taxonomic assignment.
- The consensus table gives the final assignment; the placement report shows the evidence behind it.
Screenshot guide

Resolve Taxonomy by Integrating T-BAS and Reference Databases.
Output file: taxonomy_best_consensus.csv
- Open the consensus table to view the best-supported taxonomic assignment for every ASV.
- The consensus table combines T-BAS phylogenetic placement with reference database matches to assign a single best-supported taxonomy for downstream analyses.
Screenshot guide

View the DESeq2 Relative Abundance Plot.
Output file: abundance_relative_genus_stacked.png
Relative abundance: plots show how fungal community composition varies across survey dates and treatments.
This figure shows
Treatment groups: High, low, and no treatment.
Survey dates: Each bar represents one sampling date.
Relative abundance: Bar segments show the proportion of reads assigned to each genus.
Dominant genera: Colors identify the most abundant genera.
Screenshot guide

Interpret the DESeq2 Detailed Report.
Output file: deseq2_results_OTU_detailed_analysis_report.txt
- Use the detailed report to identify which taxa differ significantly among Treatment groups.
Overall Treatment effect: 15 taxonomic features showed a significant Treatment effect by omnibus LRT (FDR ≤ 0.05).
High vs low: Neoascochyta was depleted in high relative to low.
NONE vs low: Pyrenochaeta and Pyrenochaetopsis were enriched in NONE relative to low.
Across contrasts: No taxa were consistently enriched or depleted in both pairwise contrasts.
- As you read the report, ask:
- Which taxa are significantly different between groups?
- Is each taxon enriched or depleted relative to the low reference group?
- Are any taxa consistently changed in both pairwise contrasts?
Take-home message: DESeq2 tests Treatment (high, low, NONE); Survey Date is analyzed separately for temporal community change.
Screenshot guide

Examine Faith’s Phylogenetic Diversity Results.
Output file: FaithPD_observed_OTUtree.png
- Faith’s PD summarizes phylogenetic diversity across survey dates while accounting for repeated sampling by Plot.
- Survey date significantly affected phylogenetic diversity after accounting for plot-to-plot variation.
- Overall test: Friedman test (blocked by Plot)
- Group differences: Different letters indicate significantly different groups.
- Diversity trend: Faith's PD increases over time.
Screenshot guide

Practical 1a Summary.
You have completed the workflow
- Phylogenetic placement using EPA-ng
- Placement confidence assessment
- Faith's PD, UniFrac, and Bray–Curtis/NMDS analyses
- Taxonomy resolution
- DESeq2 differential abundance
- Publication-ready figures and reports
Key take-home messages
- Phylogenetic placement provides evolutionary context for unknown sequences.
- Alpha and beta diversity reveal complementary aspects of community change.
- Taxonomy resolution and DESeq2 identify the taxa associated with those patterns.
- Integrating these analyses provides a biological interpretation of microbial community change.
Screenshot guide

In-Class Exercise 4.
- Single-Locus vs. Multilocus Placement
Before class
- Complete Practical 1A and Practical 1B.
- Be familiar with the T-BAS phylogenetic placement workflow.
- Place unknown sequences using a single locus.
- Repeat the placement using multiple loci.
- Compare taxonomic assignments and EPA-ng/gappa LWR support.
- Discuss as a class why the single- and multilocus results agree or differ.
Learning goal
- Evaluate how additional loci affect phylogenetic placement and taxonomic confidence.
- Come prepared to apply the workflow independently to a new set of unknown sequences.
Screenshot guide
