A virtual infrastructure defined through the paradigm of software as code. The infrastructure comprises a distributed storage, a processing cluster, an authentication service, a container repository and a set of template applications for accessing and processing the data. It is described in Ansible and RADL (Resource and Application description Language) and can be instantiated using Infrastructure Manager (www.grycap.upv.es/im) in a wide range of cloud backends.
This result has been implemented and demonstrated as the backend which supports the storage and processing of the Medical Imaging and Clinical data of the CHAIMELEON Repository.
This virtual infrastructure provides dataset traceability, offering a proxy service to annotate datasets, machine learning models and accesses. Our traceability solution addresses the whole cycle from dataset creation to publication of models. This provides traceability for data owners and data scientists.
CancerAddressing target audiences and expressing needs
- Business partners – SMEs, Entrepreneurs, Large Corporations
- Expanding to more markets /finding new customers
- Collaboration
We are looking for new use cases where our solution can be implemented and demonstrated. Our Processing Virtual Infrastructure is suitable for any application that requires a data analytics processing back-end based on a filesystem. Its unique functionalities for dataset traceability offer an important differentiating factor for use cases where the tracing of datasets, models and software versions is a crucial feature (e.g. reproducibility, authorship and ownership, etc.)
- Other Actors who can help us fulfil our market potential
- Research and Technology Organisations
- Academia/ Universities
R&D, Technology and Innovation aspects
This technology will be demonstrated in the cloud-based CHAIMELEON Repository, an R&D pilot infrastructure created in the framework of a H2020 RIA Project. The project includes internal and external validation phases.
The solution offered has excellent:
-Performance: with distributed and processing resources;
-Efficiency: with automatic horizontal elasticity for processing;
– Openness: components with open licenses for commercial exploitation.
The replicability of this software solution is high as it can be redeployed in the same conditions with minimal user intervention.
The sustainability of this solution is high as it enables low operational costs, since deploying and operating the platform does not require expert knowledge on cloud computing. In addition, the system uses open and widespread components with open licenses for commercial exploitation.

