Update DENV Nextclade datasets with 2026 lineage definitions - #478
jamessiqueirap wants to merge 7 commits into
Conversation
|
Hello @rneher @ivan-aksamentov @corneliusroemer, how are you guys doing? I'm requesting to merge the dengue lineages dataset. I've included a short summary of the update in the description. If you notice any issues, please let me know. Thanks! |
|
Note ✨ AI is responding on behalf of Ivan Review of PR #478, the 2026 update of TestingWhat changedChanges found in the diff, including the ones not mentioned in the description [click to expand]The PR replaces
Non-blocking issues🟡 F1. RESTRICTED Pathoplexus sequences need a data use statement [click to expand]On 2026-09-24, 36 of the Pathoplexus accessions in this PR had
In these new lineages, every tree tip is RESTRICTED, so replacing them with Open sequences is not an option:
The Pathoplexus terms allow sharing RESTRICTED data onward in web applications if the terms are clearly communicated and each sequence shows its Pathoplexus accession (Data Use Terms §4.2.4). The tip names already show the accessions. But the READMEs do not mention Pathoplexus, and Suggestions:
🟡 F2. CHANGELOG does not list the new lineages [click to expand]The new entry ( Suggestions:
🟡 F3. README says sequences may be assigned as "Outgroup", but the trees no longer use this label [click to expand]All READMEs say that sequences from other serotypes "may be assigned as 'Outgroup'" ( Suggestions:
🔵 F4. Country data was removed, but the display settings still refer to it [click to expand]You added the Suggestions:
🔵 F5. "Only publicly available data" is not quite accurate [click to expand]The CHANGELOG says that the tree sequences were updated "to retain only publicly available data". The trees still contain GISAID tips and one sequence without a public accession.
GISAID tips are also used in other datasets in this repository (for example Suggestions:
🔵 F6. Placement masks do not match the GFF3 UTR boundaries [click to expand]This was already the case before this PR, but the trees are regenerated here, so it is easy to fix now. Placement masks are 0-based and end-exclusive, so the expected masks are
The DENV-2 masks are the same as the DENV-1 masks and end after the end of the DENV-2 genome. The effect on placement is small: 1 to 4 nt next to the coding region are not masked, and Nextclade does not fail on the out-of-range end. Suggestions:
🔵 F7. The linked workflow repository does not contain the 2026 update yet [click to expand]
Suggestions:
QuestionsQuestions about intent [click to expand]
Clade distributionTip counts per lineage label. Sequences are matched by INSDC accession without the version suffix, so tips that were only renamed to Pathoplexus accessions count as unchanged. A sequence that moved to a child lineage counts as removed from the parent and added to the child. Only rows that changed are shown. DENV-1: 387 -> 417 tips [click to expand]
DENV-2: 438 -> 515 tips [click to expand]
DENV-3: 189 -> 209 tips [click to expand]
DENV-4: 149 -> 160 tips [click to expand]
Validation summaryValidation checks [click to expand]Lineages
GFF3 annotation and reference (unchanged in this PR)
Tree metadata
Data sources
Build
Nextclade CLI 3.23.0
NotesClick to expand
|
|
@jamessiqueirap Thank you for the update! As an engineer I have no technical complaints. I asked my AI to screen these changes further. If you agree with any points, please feel free to reply and/or update the dataset. I will let my human scientist colleagues to review the scientific parts. |
|
@jamessiqueirap this looks good. But I since this is using data from Pathoplexus, we need to make sure that we link back to the original source, and the DataUseTerms. For example data, it would be best if this contained only Open data. One way to do this would be to include this provenance info into the tree such that the tool tips look like this:
To this end, you need to add URL fields to the metadata. We have a script that does this for many workflows and it looks like this: https://github.com/nextstrain/rsv/blob/master/ingest/bin/curate-urls.py Then the auspice config should contain the metadata columns to use: https://github.com/nextstrain/rsv/blob/master/config/auspice_config.json#L96-L100 |
|
The nextstrain docs page has some more detailed instructions: |
|
Hi guys @ivan-aksamentov @rneher, thanks for the feedback! I’m working through all the points and will resubmit it here shortly. Thanks! |
|
Thanks @rneher and @ivan-aksamentov for the detailed review. I’ve revised the DENV update to address the points raised:
Richard, I also followed your [suggestion regarding source and data-use links](#478 (comment)): accession and data-use fields are now prepared as linked tree-tip metadata. I’m also considering changing the institution name in the directory structure. Would there be any issue with doing that? In any case, considering the addition of the new lineages included in this update, I was thinking of releasing this functional update first and making those structural changes afterwards. What do you think? |
|
@jamessiqueirap thanks so much for updating the datasets and bringing them inline with the data use terms. We try to have an eye on this to maintain the trust people have in the data sharing mechanism on pathoplexus regarding the name/directory change... this is in principle possible. But it might break links to the datasets that other people have bookmarked or hard-coded. It might be possible to include the old links as "short-cuts" into the pathogen json. But I am not sure whether such short cuts are allowed for names that were previously full path. @ivan-aksamentov would know more. One could also envision to move the dengue datasets into a top-level group like the nextstrain group. e.g. |
|
Thanks, @rneher ! That makes sense, and preserving existing links is definitely something I’d want to avoid breaking. |

Description of proposed changes
This annual update incorporates 23 new lineages defined by the scientific committee across the four DENV serotypes, increasing the total from 218 to 241.
Test and representative sequences have also been updated to support the continued use of publicly available data. Test and representative datasets are now composed primarily of sequences available from Pathoplexus.
Checklist