Skip to main content

Practical 1a. Phylogeny-Based Analysis of DNA Barcode Sequences

This tutorial provides step-by-step instructions for phylogenetic placement and analysis using T-BAS and DeCIFR.


Part 1 — Taxonomy Resolution with Multiple Databases

Integrate phylogenetic placement with multiple reference databases to improve taxonomic resolution and evaluate assignment support.

Task 1

Why Phylogenetic Placement?

Why it matters

  • Places unknown ITS sequences into an evolutionary context.
  • More informative than similarity searches alone.
  • Estimates placement confidence.
  • Supports downstream diversity analyses.

Key outputs

  • Evolutionary placement
  • Likelihood weights
  • Interactive tree

Goal: Determine where unknown sequences belong on the fungal tree of life.

Screenshot guide
Why Phylogenetic Placement?
Task 2

Step 1. Select a Reference Tree.

What you will do

  • Select the Fungi v3 reference tree.
  • Download the example datasets.
  • Prepare the ITS ASV FASTA file.

Observe

  • Choose the reference matching the ITS marker.
  • The tutorial uses Fungi v3.
  • Next: Upload the ITS sequences.
Screenshot guide
Step 1. Select a Reference Tree
Task 3

Practical 1a. Phylogenetic Placement of DNA Barcode Sequences.

  • Click here to select reference tree
Screenshot guide
Practical 1a. Phylogenetic Placement of DNA Barcode Sequences
Task 4

Select Fungi v3 for phylogenetic placement of ITS reads.

  • Target reference tree for placement
Screenshot guide
Select Fungi v3 for phylogenetic placement of ITS reads
Task 5

Download example files.

  • Click on examples
Screenshot guide
Download example files
Task 6

Download the metadata and reference paper used in this tutorial. A ZIP archive containing all files is also available.

Screenshot guide
Download the example2 datasets, metadata, and reference paper used in this tutorial. A ZIP archive containing all files is also available.
Task 7

Step 2. Upload ITS Sequences.

What you will do

  • Upload the ASV FASTA file.
  • Enable ITS filtering.
  • Select RDP, FUNGuild and NCBI RefSeq.

Common mistakes

  • Wrong FASTA format
  • Mixed non-ITS sequences
  • Wrong reference tree
  • ITSx identifies non-ITS sequences before placement.
Screenshot guide
Step 2. Upload ITS Sequences
Task 8

Drag the ASV fasta file into the unknown query box.

Screenshot guide
Drag the ASV fasta file into the unknown query box
Task 9

Select filter unknowns and generate UNITE report.

Screenshot guide
Select filter unknowns and generate UNITE report
Task 10

Select RDP Classifier, FUNGuild, and NCBI ITS RefSeq databases.

Screenshot guide
Select RDP Classifier, FUNGuild, and NCBI ITS RefSeq databases
Task 11

Step 3. Configure EPA-ng.

What you will do

  • Choose EPA-ng.
  • Provide a run name.
  • Select the ITS locus.
  • Submit the analysis.

Runtime

  • ~1 hour
  • Monitor progress
  • Email notification
  • EPA-ng is the longest-running step in the tutorial.
Screenshot guide
Step 3. Configure EPA-ng
Task 12

Use EPA-ng (Evolutionary Placement Algorithm – Next Generation).

Screenshot guide
Use EPA-ng (Evolutionary Placement Algorithm – Next Generation)
Task 13

Provide a label for the run for future reference.

Screenshot guide
Provide a label for the run for future reference
Task 14

Select the ITS locus in the pull-down menu.

Screenshot guide
Select the ITS locus in the pull-down menu
Task 15

Step 4. Review the ITSx Report.

What you will do

  • Review the ITSx quality-control report.
  • Identify ITS-only sequences.
  • Trim or remove sequences, if necessary, before placement.

Observe

  • Most sequences should contain the expected ITS region.
  • Problematic sequences can be excluded before analysis.
  • Quality control improves the accuracy of downstream phylogenetic placement.
Screenshot guide
Step 4. Review the ITSx Report
Task 16

Verify ITS-Only Sequences using ITSx.

Screenshot guide
Verify ITS-Only Sequences using ITSx
Task 17

ITSx report reveals ITS1-only sequences.

Screenshot guide
ITSx report reveals ITS1-only sequences
Task 18

Have the option to submit, cancel or trim sequences.

  • Because these are ITS1-only sequences, click Submit to continue
Screenshot guide
Have the option to submit, cancel or trim sequences
Task 19

Step 5. Monitor the EPA-ng Analysis.

What you will do

  • Monitor the progress bar.
  • Wait for the analysis to complete (~1 hour).
  • Open the completed run directory.
  • View the summary report.

Expected outputs

  • Run summary
  • Placement statistics
  • Interactive tree
  • EPA-ng sends an email when the analysis is complete.
Screenshot guide
Step 5. Monitor the EPA-ng Analysis
Task 20

Progress bar displays status of run.

  • Takes about an hour to finish
Screenshot guide
Progress bar displays status of run
Task 21

Summary reports at end of run and option to view tree.

  • Click to view tree
Screenshot guide
Summary reports at end of run and option to view tree
Task 22

Completed run is also announced via email.

Screenshot guide
Completed run is also announced via email
Task 23

Step 6. Explore the Interactive Tree.

What you will do

  • Open the interactive phylogenetic tree.
  • Display taxonomic groups.
  • Locate the placed query sequences.
  • Inspect nearby reference taxa.

Observe

  • Unknown sequences are placed within an evolutionary context.
  • Closely related taxa provide biological interpretation.
  • Phylogenetic placement is more informative than a simple database match.
Screenshot guide
Step 6. Explore the Interactive Tree
Task 24

Tree view in T-BAS showing Class and Unknowns shaded in grey.

Screenshot guide
Tree view in T-BAS showing Class and Unknowns shaded in grey
Task 25

Step 7. Evaluate Placement Confidence.

What you will do

  • Color branches using EPA-ng LWR support.
  • Compare EPA-ng placement support with gappa taxonomic assignment support.
  • Inspect individual placements.

Interpretation

  • High LWR = confident placement.
  • Low LWR may indicate ambiguous taxonomy or missing references.
  • Confidence values are essential for interpreting placement results.
Screenshot guide
Step 7. Evaluate Placement Confidence
Task 26

Display EPA-ng and gappa LWR Support.

  • Select EPA and gappa
Screenshot guide
Display EPA-ng and gappa LWR Support
Task 27

Interpret EPA-ng and gappa LWR Support.

  • High placement support does not necessarily mean high-confidence species identification.

EPA-ng LWR

  • Placement support
  • Where does the query place?

gappa LWR

  • Taxonomic assignment support
  • Do the plausible placements support the same taxonomic assignment?

Interpret the two values together

  • High EPA + high gappa: Confident placement and confident taxonomic assignment.
  • High EPA + low gappa: Placement may be strong, but species-level assignment remains uncertain.

Take-home message: Strong phylogenetic placement ≠ automatically strong species assignment.

Screenshot guide
Interpret EPA-ng and gappa LWR Support
Task 28

Step 8. Inspect Sequence Alignments.

What you will do

  • Select a focal clade.
  • Open the multiple sequence alignment viewer.
  • Compare query sequences with nearby references.

Observe

  • Sequence similarity
  • Conserved positions
  • Potential alignment problems
  • Alignment inspection helps validate unusual placements.
Screenshot guide
Step 8. Inspect Sequence Alignments
Task 29

Select the Focal Clade.

  • Uncheck Use branch length.
  • Set Circle diameter multiplier = 20.
  • Locate the focal clade.
  • Click the node defining the clade to select it.
  • Next: zoom in to make the focal clade easier to select.
Screenshot guide
Select the Focal Clade
Task 30

Zoom In to Select the Focal Clade.

  • Use + to zoom into the focal region.
  • Locate the branch defining the clade of interest.
  • Click the node defining the clade to select it.
Screenshot guide
Zoom In to Select the Focal Clade
Task 31

View the Alignment for the Selected Clade.

  • Click View to open the alignment
Screenshot guide
View the Alignment for the Selected Clade
Task 32

Compare query sequences with nearby references.

  • Look for:
  • Sequence similarity
  • Conserved and variable sites
  • Unexpected gaps
Screenshot guide
Compare query sequences with nearby references
Task 33

In-Class Discussion 1.

  • Interpreting Phylogenetic Placement

Before coming to class, think about

  • Where did most of your ITS sequences place on the reference tree?
  • Which placements had high confidence? Which were uncertain?
  • Why might some sequences have ambiguous placements?
  • What advantages does phylogenetic placement provide over a BLAST similarity search?

Learning goal

  • Explain how phylogenetic placement provides an evolutionary framework for identifying unknown DNA barcode sequences.
  • Complete Practical 1A before class and bring your results for discussion.
Screenshot guide
In-Class Discussion 1

Part 2 — Microbial Community Diversity

Analyze phylogenetic and community diversity using complementary alpha- and beta-diversity approaches.

Task 34

Why Analyze Microbial Community Diversity?

Why analyze diversity?

  • Compare microbial communities using complementary phylogenetic and abundance-based approaches.
  • Measure diversity within and differences among samples.
  • Quantify changes across treatments and sampling dates.
  • Link phylogeny with ecological interpretation.

Two complementary views of diversity

Alpha diversity — within a sample

  • Faith's PD — phylogeny-based diversity
  • How much evolutionary diversity is present?

Beta diversity — among samples

  • UniFrac — phylogeny-based dissimilarity
  • Bray–Curtis — abundance-based dissimilarity
  • NMDS (Non-metric Multidimensional Scaling) — visualizes Bray–Curtis dissimilarities
  • How different are microbial communities?

Take-home: Alpha diversity describes diversity within samples; beta diversity describes differences in community composition among samples.

Screenshot guide
Why Analyze Microbial Community Diversity?
Task 35

Step 9. Launch UniFrac Analysis.

What you will do

  • Open the UniFrac analysis.
  • Upload the ASV count table.
  • Upload the sample metadata.
  • Submit the analysis.

Runtime

  • ~1 minute
  • Results stored in the run directory.
  • UniFrac compares communities using branch lengths on the reference tree.
Screenshot guide
Step 9. Launch UniFrac Analysis
Task 36

Perform diversity analysis using UniFrac.

  • Click on UniFrac
Screenshot guide
Perform diversity analysis using UniFrac
Task 37

Upload the ASV counts file.

Screenshot guide
Upload the ASV counts file
Task 38

Upload the sample metadata file with experimental attributes.

  • Include a note about the run and click submit
Screenshot guide
Upload the sample metadata file with experimental attributes
Task 39

Wait for the run to complete.

  • Takes a minute to finish
Screenshot guide
Wait for the run to complete
Task 40

Step 10. Review UniFrac Results.

What you will do

  • Open the completed run directory.
  • Locate the Faith_pd and matplotlib folders.
  • Review tables and publication-ready figures.

Key outputs

  • Faith_pd/: statistics
  • matplotlib/: figures
  • Most analysis outputs are organized automatically.
Screenshot guide
Step 10. Review UniFrac Results
Task 41

Click on the run directory to view the results.

Screenshot guide
Click on the run directory to view the results
Task 42

Step 11. Interpret Faith's PD.

What you will do

  • Open the Faith’s PD figure.
  • Compare phylogenetic diversity across sampling dates.
  • Examine the overall temporal pattern.

Observe

  • Faith’s PD generally increases across the 2022 sampling dates.
  • Later samples tend to have greater phylogenetic diversity.
  • Faith's PD measures the total evolutionary diversity within each sample.
Screenshot guide
Step 11. Interpret Faith's PD
Task 43

Locate the Faith_pd/ Output Folder.

  • Faith's PD results and figures are saved in Faith_pd/. Open this folder to examine phylogenetic alpha diversity.
  • Click Faith_pd/Faith's PD resultsand figures
Screenshot guide
Locate the Faith_pd/ Output Folder
Task 44

Visualize Faith's Phylogenetic Diversity.

Screenshot guide
Visualize Faith's Phylogenetic Diversity
Task 45

In-Class Discussion 2.

  • Interpreting Phylogenetic Diversity

Before coming to class, think about

  • Which sampling dates have the greatest phylogenetic diversity?
  • Do the results agree with your expectations?
  • What biological processes could explain the observed pattern?
  • Why might Faith's PD differ from species richness?
  • Complete Practical 1A before class and bring your results for discussion.

Learning goal

  • Interpret changes in phylogenetic diversity and relate them to microbial community dynamics.
Screenshot guide
In-Class Discussion 2
Task 46

Step 12. Interpret NMDS.

What you will do

  • Open the Bray–Curtis NMDS plot.
  • Compare clustering by survey date.
  • Relate ordination to Faith's PD.

Observe

  • Samples separate over time.
  • Community composition shifts through the season.
  • Ordination complements diversity metrics by revealing community structure.
Screenshot guide
Step 12. Interpret NMDS
Task 47

Locate the matplotlib/ Output Folder.

  • Diversity results are organized into several output folders. Open matplotlib/ to view publication-ready figures with statistical annotations.
  • Click matplotlib/Publication-ready figures with statistical annotations.
Screenshot guide
Locate the matplotlib/ Output Folder
Task 48

Visualize Temporal Changes in Microbial Community Composition.

Screenshot guide
Visualize Temporal Changes in Microbial Community Composition
Task 49

In-Class Discussion 3.

  • Comparing Microbial Communities

Before coming to class, think about

  • Which samples cluster together?
  • Which sampling dates are most distinct?
  • Do the NMDS and Faith's PD results tell the same story?
  • What ecological processes might explain these community shifts?
  • Complete Practical 1A before class and bring your results for discussion.

Learning goal

  • Interpret differences in community composition using beta-diversity analyses.
Screenshot guide
In-Class Discussion 3

Part 3 — Taxonomy Resolution with Multiple Databases

Integrate phylogenetic placement with multiple reference databases to improve taxonomic resolution and evaluate assignment support.

Task 50

Why Taxonomy Resolution Is Needed.

  • Phylogenetic placement accurately places sequences on a reference tree, but resolution is limited by the taxa represented in that tree.
  • Reference databases (UNITE, RDP, and NCBI RefSeq) provide broader coverage but can contain conflicting or incomplete classifications.
  • The Taxonomy Resolver integrates T-BAS placements with multiple databases to recover the best-supported taxonomy.

Take-home message: Integrating phylogenetic placement with multiple reference databases produces more robust taxonomic assignments than either approach alone.

Screenshot guide
Why Taxonomy Resolution Is Needed
Task 51

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Launch the Taxonomy Resolver
  • Integrate T-BAS placements with UNITE, RDP, and NCBI RefSeq to recover the best-supported taxonomy beyond the limits of any single reference tree.
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 52

Step 13. Configure the Taxonomy Resolver.

What you will do

  • Enter the T-BAS accession.
  • Upload the ASV count table.
  • Upload the metadata table.
  • Choose the recommended taxonomy strategy.

Required inputs

  • Placement accession
  • ASV counts
  • Metadata
  • Read mode
  • These inputs provide the information required for downstream analyses.
Screenshot guide
Step 13. Configure the Taxonomy Resolver
Task 53

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Enter T-BAS accession
  • Upload ASV count table
  • Specify read mode
  • Upload metadata table
  • Choose the recommended taxonomy strategy
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 54

Step 14. Enable Downstream Analyses.

What you will do

  • Enable DESeq2.
  • Enable Faith's PD.
  • Complete all required fields.
  • Submit the analysis.

Runtime

  • ~1 minute
  • Results saved under analysis_output/
  • The Taxonomy Resolver can automatically launch downstream analyses.
Screenshot guide
Step 14. Enable Downstream Analyses
Task 55

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Check “Run” to enable DESeq2 analysis
  • Click the warning banner to step through the required fields.
  • Check “Run” to enable Faith's PD analysis
  • Select the downstream analyses to perform after taxonomy resolution.
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 56

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Complete the required fields
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 57

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Complete the three required fields (red *). For repeated-measures studies, also specify the block column (blue *). Click the warning banner to reveal any remaining required fields.
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 58

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Complete the highlighted required fields
  • Continue clicking the warning banner until all required fields have been completed.
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 59

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Confirm the required grouping fields:
  • Treatment and Survey_Date
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 60

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • Enter a run name, then click Extract taxonomy
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 61

Step 15. Explore the Results.

What you will do

  • Open the consensus taxonomy table.
  • Review the best-supported taxonomy.
  • Locate DESeq2 and Faith's PD output folders.
  • Inspect publication-ready figures.

Key outputs

  • Consensus taxonomy
  • DESeq2 figures
  • Faith's PD figures
  • Downstream Taxonomy Resolver outputs are written to analysis_output/; placement/support reports remain in the run directory.
Screenshot guide
Step 15. Explore the Results
Task 62

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

  • The analysis typically completes in ~1 minute
  • Open the run directory to view the results
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 63

Example Output Paths.

The placement/support table is in the run directory; downstream Taxonomy Resolver results are under analysis_output/.

Screenshot guide
Example Output Paths
Task 64

Step 16. Interpret the Results.

  • Locate the placement-support report generated by the Taxonomy Resolver.
  • Interpret EPA-ng and gappa likelihood weight ratio (LWR) values.
  • Review the consensus taxonomic assignment.
  • Interpret downstream diversity and differential-abundance results.

Questions to ask

  • Is the phylogenetic placement well supported?
  • Is the taxonomic assignment well supported?
  • Do placement and taxonomic support agree?
Screenshot guide
Step 16. Interpret the Results
Task 65

Examine Placement and Taxonomic Support.

Output file: assignments_report_pretty_withgappa(T-BAS run ID).csv

In this example: assignments_report_pretty_withgappa6FAEQW0E.csv

  • Examine the placement report to evaluate placement and taxonomic support for each ASV.
Screenshot guide
Examine Placement and Taxonomic Support
Task 66

Interpret the Four Placement-Support Columns.

EPA Likelihood Weight Ratio (LWR)

Relative support among alternative phylogenetic placements of the query.

gappa Likelihood Weight Ratio (LWR)

Likelihood weight associated with the taxonomic assignment.

gappa fraction of placements (fract)

Fraction of candidate placements supporting that taxonomic assignment.

gappa accumulated LWR (aLWR)

Accumulated likelihood weight supporting that taxonomic assignment.

Key distinction: EPA-ng evaluates where the sequence can be placed on the reference tree; gappa summarizes how those placements support taxonomy.

Screenshot guide
Interpret the Four Placement-Support Columns
Task 67

How to Read EPA-ng and gappa Support.

  • EPA-ng LWR near 1.0: placement support is concentrated on one branch.
  • Split EPA-ng LWR values: multiple placements remain plausible.
  • High gappa support: the placements consistently support the same taxonomic assignment.
  • Low gappa support: taxonomic assignment remains uncertain even when a placement can be made.
  • Interpret both together:High EPA-ng + high gappa support → confident placement and taxonomySplit/low support → treat the assignment cautiously
Screenshot guide
How to Read EPA-ng and gappa Support
Task 68

From Placement Support to Consensus Taxonomy.

  • Use EPA-ng LWR to assess confidence in phylogenetic placement.
  • Use gappa support to assess confidence in the associated taxonomy.
  • Use the Taxonomy Resolver consensus table for the final taxonomic assignment.
  • The consensus table gives the final assignment; the placement report shows the evidence behind it.
Screenshot guide
From Placement Support to Consensus Taxonomy
Task 69

Resolve Taxonomy by Integrating T-BAS and Reference Databases.

Output file: taxonomy_best_consensus.csv

  • Open the consensus table to view the best-supported taxonomic assignment for every ASV.
  • The consensus table combines T-BAS phylogenetic placement with reference database matches to assign a single best-supported taxonomy for downstream analyses.
Screenshot guide
Resolve Taxonomy by Integrating T-BAS and Reference Databases
Task 70

View the DESeq2 Relative Abundance Plot.

Output file: abundance_relative_genus_stacked.png

Relative abundance: plots show how fungal community composition varies across survey dates and treatments.

This figure shows

Treatment groups: High, low, and no treatment.

Survey dates: Each bar represents one sampling date.

Relative abundance: Bar segments show the proportion of reads assigned to each genus.

Dominant genera: Colors identify the most abundant genera.

Screenshot guide
View the DESeq2 Relative Abundance Plot
Task 71

Interpret the DESeq2 Detailed Report.

Output file: deseq2_results_OTU_detailed_analysis_report.txt

  • Use the detailed report to identify which taxa differ significantly among Treatment groups.

Overall Treatment effect: 15 taxonomic features showed a significant Treatment effect by omnibus LRT (FDR ≤ 0.05).

High vs low: Neoascochyta was depleted in high relative to low.

NONE vs low: Pyrenochaeta and Pyrenochaetopsis were enriched in NONE relative to low.

Across contrasts: No taxa were consistently enriched or depleted in both pairwise contrasts.

  • As you read the report, ask:
  • Which taxa are significantly different between groups?
  • Is each taxon enriched or depleted relative to the low reference group?
  • Are any taxa consistently changed in both pairwise contrasts?

Take-home message: DESeq2 tests Treatment (high, low, NONE); Survey Date is analyzed separately for temporal community change.

Screenshot guide
Interpret the DESeq2 Detailed Report
Task 72

Examine Faith’s Phylogenetic Diversity Results.

Output file: FaithPD_observed_OTUtree.png

  • Faith’s PD summarizes phylogenetic diversity across survey dates while accounting for repeated sampling by Plot.
  • Survey date significantly affected phylogenetic diversity after accounting for plot-to-plot variation.
  • Overall test: Friedman test (blocked by Plot)
  • Group differences: Different letters indicate significantly different groups.
  • Diversity trend: Faith's PD increases over time.
Screenshot guide
Examine Faith’s Phylogenetic Diversity Results
Task 73

Practical 1a Summary.

You have completed the workflow

  • Phylogenetic placement using EPA-ng
  • Placement confidence assessment
  • Faith's PD, UniFrac, and Bray–Curtis/NMDS analyses
  • Taxonomy resolution
  • DESeq2 differential abundance
  • Publication-ready figures and reports

Key take-home messages

  • Phylogenetic placement provides evolutionary context for unknown sequences.
  • Alpha and beta diversity reveal complementary aspects of community change.
  • Taxonomy resolution and DESeq2 identify the taxa associated with those patterns.
  • Integrating these analyses provides a biological interpretation of microbial community change.
Screenshot guide
Practical 1a Summary
Task 74

In-Class Exercise 4.

  • Single-Locus vs. Multilocus Placement

Before class

  • Complete Practical 1A and Practical 1B.
  • Be familiar with the T-BAS phylogenetic placement workflow.
  • Place unknown sequences using a single locus.
  • Repeat the placement using multiple loci.
  • Compare taxonomic assignments and EPA-ng/gappa LWR support.
  • Discuss as a class why the single- and multilocus results agree or differ.

Learning goal

  • Evaluate how additional loci affect phylogenetic placement and taxonomic confidence.
  • Come prepared to apply the workflow independently to a new set of unknown sequences.
Screenshot guide
In-Class Exercise 4