Plankton genomic demonstrator focus on two key objectives:
– Notebook 1: Exploring genetic data & identifying clusters containing unknown genes Discovery of as yet undescribed biodiversity from genetic and morphological signals from the characterisation of their geographical distributions, co-occurrences/exclusions and correlation with environmental contexts.
-Notebook 2: Mapping the geographic distribution of plankton functional gene clusters using habitat prediction models Exploration of genetic and morphological markers of plankton diversity and abundance, in particular the new ones discovered above, to predict their spatiotemporal distribution and serve as high-resolution EOVs for biological processes.
The initial users of the plankton genomics demonstrator are, primarily, scientific researchers, including taxonomists, computational ecologists and bioinformaticians with extensive knowledge of the data collected during the Tara Oceans Expedition. The end-users include scientists from plankton biogeography, marine biogeochemistry, ecosystem health, and climate science.
Oceans, seas and watersAddressing target audiences and expressing needs
- Use of research Infrastructure
The plankton genomics descriptors exploit the metagenomic data produced through the Tara Oceans campaign. Two products are derived: (1) a sequence similarity network, with clusters of genes, in which each gene sequence is tagged with functional (what enzyme it codes for) and taxonomic (what organism it comes from) labels or is left unknown if those labels cannot be assigned + tools to explore it, (2) maps of the potential worldwide distribution of the clusters or genes described above, for a given function computed using a machine learning algorithm + an easy to use environment to produce those maps.
- International Organisations (ex. OECD, FAO, UN, etc.)
- Research and Technology Organisations
- Academia/ Universities
R&D, Technology and Innovation aspects
The process to explore or reproduce them is deployed in D4science and depends on the continuous operation and potential updates of these environments; minor modifications will be operated if needed. More “customers” are coming to the platform interested in the easy-to-digest results (the maps). We explored a very small % of the existing data and more data is coming so there is a huge potential in scalability of application, to more data. Further developments are planned in the new Blue-Cloud 2026 project
Only in theory
Full proof of replicability would be a completely independent user who would re-do everything. This is unlikely to happen since new users will want to explore *other* things, functions that we did *not* explore.
Funding is secured till middle of 2026. The data products (network+cluster and maps of one target function –carbon fixation) will be published in a way that is sustainable over the long term. The process to explore or reproduce them depends on the continuous operation and potential updates of the D4Science environment. The operation of this depends on further funding of the D4science infra to scale to the proper size if there were to be many users. Minor modifications will be operated if needed. Additional EU funding will allow evolution of the VLab into a Workbench, which is aimed at replacing and increasing the capacity and performance of the current habitat modelling workflow. NB: The products are not marketable so there will not be more revenue if there are more users.
- Global

