A dataset for studying how 3D object distortions affect object recognition and visual security.
The transformations include mesh simplification and quantization, texture resizing, compression and blurring, as well as encryption and obscuration. The mesh “data hiding” profile currently uses an encryption-based proxy.
The assets are provided as three separate ZIP archives, split into volumes of at most 4 GiB. All three archives have passed 7-Zip integrity checks and file-inventory verification.
NEXTLIFE_20260930_source.zip.001NEXTLIFE_20260930_distorted.zip.001NEXTLIFE_20260930_distorted.zip.002NEXTLIFE_20260930_distorted.zip.003NEXTLIFE_20260930_distorted.zip.004NEXTLIFE_20260930_scenes.zip.001Download every volume and keep all original filenames. Place the volumes together in one download folder. Use 7-Zip and start extraction from .zip.001 only: the remaining volumes are read automatically. Do not extract the parts separately or rename them to .zip.
The download also includes:
README_EXTRACT.txt — installation and extraction instructions.dataset_info.json — archive inventory, volume sizes and checksums.SHA256SUMS.txt — SHA-256 checksums for checking the downloaded volumes.After extraction, the assets follow the repository folder structure:
Objects/
├── Originals/ Reference objects, materials and textures
└── Distorted/
├── MeshVariants/ Transformed meshes
├── TextureVariants/ Transformed textures
├── CombinedVariants/ Manifests linking meshes and textures
└── dataset_info.json Validated selection and dataset identifier
Scenes/ Five active experimental environmentsKeep the complete folder structure unchanged. A variant manifest references other files; it is not a standalone 3D object. Excluded reference objects are not included in this distribution.
Objects/Distorted/dataset_info.json preserves the validated object selection and the stable dataset identifier. It is different from the dataset_info.json supplied beside the archives, which describes the download itself.
The 3D assets are distributed separately from the code. The viewers, scripts, configuration files and metadata needed to run the experiments are available in the NEXTLIFE GitHub repository.
The tools allow users to explore objects, compare distorted variants and run an object recognition and visual security experiment. The distortions are already generated: no regeneration is required to use this download.
Install Git, Python 3.12 and 7-Zip. Open PowerShell in a working directory and run these commands, one line at a time:
git clone https://github.com/Kaldrass/Dataset_NEXTLIFE.git
cd Dataset_NEXTLIFE
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -r requirements.txtDownload all six archive volumes into one folder, for example D:\Downloads\NEXTLIFE. From the root of the cloned repository, run the following commands, adapting the download path:
$assets = "D:\Downloads\NEXTLIFE"
& "$env:ProgramFiles\7-Zip\7z.exe" x "$assets\NEXTLIFE_20260930_source.zip.001" -o.
& "$env:ProgramFiles\7-Zip\7z.exe" x "$assets\NEXTLIFE_20260930_distorted.zip.001" -o.
& "$env:ProgramFiles\7-Zip\7z.exe" x "$assets\NEXTLIFE_20260930_scenes.zip.001" -o.These commands assume that 7-Zip is installed in its standard Windows location. Alternatively, open each .zip.001 file in the 7-Zip application, choose Extract, and select the root of the cloned repository as the destination.
Extract into a fresh clone to avoid mixing asset versions. The archives create Objects/Originals/, Objects/Distorted/ and Scenes/ directly beside metadata.json and ExperimentSecurity/. Do not extract into an additional Objects/ folder.
Allow enough disk space for both the downloaded archives and approximately 50 GiB of extracted assets.
Calculate the SHA-256 checksum of each volume and compare it with the corresponding entry in SHA256SUMS.txt. For example:
Get-FileHash "$assets\NEXTLIFE_20260930_source.zip.001" -Algorithm SHA256Repeat this check for the other five volumes.
With the 7-Zip command-line tool installed, extract from the repository root using the following commands. Adapt the download path:
7zz x /path/to/downloads/NEXTLIFE_20260930_source.zip.001 -o.
7zz x /path/to/downloads/NEXTLIFE_20260930_distorted.zip.001 -o.
7zz x /path/to/downloads/NEXTLIFE_20260930_scenes.zip.001 -o.For the Python commands in this guide, create the environment with python3 -m venv .venv and use .venv/bin/python instead of .\.venv\Scripts\python.exe. Use forward slashes in script paths.
From the repository root:
.\.venv\Scripts\python.exe ExperimentSecurity\build_dataset_catalog.pyThis command prepares a demonstration with up to 20 trials and 5 distinct objects:
.\.venv\Scripts\python.exe ExperimentSecurity\build_recognition_experiment.py --max-objects 5 --max-trials 20 --max-faces 0 --choices 6 --seed 20260622This replaces recognition_trials.json. Back up any existing trial set before changing it. The demonstration checks that the tools work; it is not a balanced experimental design.
.\.venv\Scripts\python.exe -m http.server 8015 --bind 127.0.0.1Keep the terminal open, then open one of these addresses in Chrome:
These links work on the computer running the local server. The viewers load Three.js from the Internet. If the scenes are missing, a fallback environment is used: obtain the intended scenes before running an actual study. To stop the server, press Ctrl + C in the terminal.
Responses are stored in the browser and are not automatically sent to the server.
Analysis commands and detailed usage instructions are available in the ExperimentSecurity documentation on GitHub.
ExperimentDSIS/ contains a separate historical experiment. The Old/ archives, temporary files and earlier generations are not needed to use these assets.
These instances of the FrigoTruck database were generated and polluted with the library SODGAUP (https://github.com/mathildemarcy/SODGAUP/).
All instances include the same levels of artificial unicity. Relations sizes for each instance are presented in the table below, along with the level of artificial unicity.
| relation | AU level | 500 | 5,000 | 10,000 | 25,000 | 50,000 | 100,000 |
|---|---|---|---|---|---|---|---|
| Address | - | 4,366 | 44,068 | 88,391 | 220,256 | 438,412 | 879,248 |
| Brand | 0.85 | 488 | 488 | 488 | 488 | 488 | 488 |
| Chassis | 1 | 4,366 | 44,068 | 88,391 | 220,256 | 438,412 | 879,248 |
| Cell | 1 | 4,366 | 44,068 | 88,391 | 220,256 | 438,412 | 879,248 |
| Compressor | 1 | 6,782 | 68,031 | 136,576 | 340,146 | 676,698 | 1,360,058 |
| Equipment_category | 0 | 5 | 5 | 5 | 5 | 5 | 5 |
| Evaporator | 1 | 6,889 | 71,574 | 143,935 | 359,052 | 712,502 | 1,430,955 |
| Event | - | 4,366 | 44,068 | 88,391 | 220,256 | 438,412 | 879,248 |
| Event_type | 0 | 8 | 8 | 8 | 8 | 8 | 8 |
| Manufacturer | 0.85 | 125 | 125 | 125 | 125 | 125 | 125 |
| Model | 0.85 | 1000 | 1000 | 1000 | 1000 | 1000 | 1000 |
| Occupation | 0.6 | 8 | 8 | 8 | 8 | 8 | 8 |
| Owner | - | 4,366 | 44,068 | 88,391 | 220,256 | 438,412 | 879,248 |
| Refrigerant | 0.75 | 26 | 26 | 26 | 26 | 26 | 26 |
| Region | 0.7 | 8 | 8 | 8 | 8 | 8 | 8 |
| TRS | 1 | 4,366 | 44,068 | 88,391 | 220,256 | 438,412 | 879,248 |
| TRS_category | 0.4 | 7 | 7 | 7 | 7 | 7 | 7 |
| TRU | 1 | 4,366 | 44,068 | 88,391 | 220,256 | 438,412 | 879,248 |
More information on these databases and their generation and pollution is available at https://github.com/mathildemarcy/SODGAUP/.
The 6DoF Physiological Dataset is a publicly available collection of electrodermal activity (EDA) and electrocardiogram (ECG) recordings acquired during a 6DoF virtual reality experiment. Participants performed a series of predefined movements to investigate the effects of user motion and sensor placement on physiological signal quality.
The dataset was used in our study, "Signal or Noise? How User Movement and Sensor Placement Impact Physiological Data Collection in 6DoF Virtual Reality."
It includes 430 physiological recordings, categorized according to the sensor placement on the body, and is distributed in CSV format.
The GuidedSAM-Plume dataset targets the segmentation of volcanic ash plumes and clouds in real-world observational imagery.
It comprises 47 annotated time-series sequences, totaling 759 frames, collected using ground-based camera systems during field campaigns and during the monitoring of explosive volcanic eruptions. The sequences have a mean length of 16 frames and an average spatial resolution of 2996 × 1685 pixels.
The dataset covers a broad range of acquisition and environmental conditions, including varying viewing angles, illumination, occlusions, and weather, as well as diverse plume dynamics from early formation to fully developed structures.
Two annotation sources are provided, both produced by expert volcanologists: manual segmentations and segmentation masks generated using the GuidedSAM-Plume interactive annotation framework.
The SeracFallDet dataset is a specialized sub-task of change detection focused on identifying volumetric changes. Specifically, this dataset contains images with serac fall (glacier ice collapse) annotations.
It consists of annotated pairs of high resolution images captured using fixed terrestrial time-lapse cameras. Detecting serac falls is critical for natural hazard monitoring, enabling better prevention of catastrophic events. Each annotation are masks generated by bounding boxes fusion.
Composition.
The dataset includes 11 scenes from glaciers (currently only within the French Alps), comprising:
- 8199 pre-registered images
- 2123 annotations of serac falls, represented as polygonal annotations (derived from bounding boxes) around detected changes, if any.
Acknowledgements.
We extend our sincere gratitude to all following contributors for providing access to the glacier footage that forms the dataset: Voltalia, DDT-74 and IGE, communauté de commune de la Vallée de Chamonix Mont-Blanc (CCVCMB) and COMPAGNIE DU MONT BLANC/Aiguille du Midi www.montblancnaturalresort.com).
This dataset provides rendered views, binary masks, and patch data for 5 mesh and point cloud datasets (TMQ, TSMD, SJTU-TMQA, BASICS and WPC) that are suited to be used with Graphics-LPIPS-QualCompare. They have been generated using QualCompare.
This repository accompanies the paper:
Rendering Matters: Reproducible Image-based Quality Assessment of 3D Meshes and Point Clouds with QualCompare
Guillaume Lavoué, Gautier Campagne, Florent Dupont
Currently under review.
The Multi-View Evaluation Protocol for Glass Container Inspection (MVEP) is a specialized benchmark dataset designed to evaluate multi-view fusion methods for industrial quality control of transparent materials. MVEP addresses the critical challenge of automated defect detection and severity assessment on glass container surfaces, where view-dependent optical phenomena (including specular reflections, refractions, and transparency effects) significantly limit the reliability of single-view inspection. The dataset comprises 16,000 synchronized multi-view images of glass containers captured from six calibrated viewpoints under controlled industrial lighting conditions. Each container is annotated with object-level bounding boxes and ordinal severity labels corresponding to surface degradation defects, ranging from minimal visual alteration to critical damage requiring rejection. The defect taxonomy focuses on erasure-type surface defects, replacing earlier wipe-based definitions to better reflect realistic industrial degradation patterns, and excludes any reference to scuffing. MVEP provides a realistic distribution of quality grades encountered in production environments and includes inherent annotation uncertainty due to the subjective nature of visual quality assessment. The dataset is particularly well suited for evaluating ordinal classification methods, multi-view fusion strategies, cross-view consistency constraints, and robustness to annotation noise in industrial inspection scenarios involving transparent materials.
The Multi-Modal and Multi-View Object Detection Dataset (MMDOD) is a comprehensive benchmark designed to advance research in detection-driven image fusion under strong modality-view dependencies. MMDOD contains over 10,000 high-resolution images of transparent glass containers captured under four complementary imaging modalities (visible light, near-infrared (NIR), low-contrast, and polarization shift) across six distinct viewpoints. Each image is annotated with detailed object-level bounding boxes and class labels, enabling rigorous evaluation of multi-modal and multi-view fusion methods for object detection tasks. The dataset addresses a critical gap in existing benchmarks by providing synchronized multi-modal multi-view observations of transparent materials, which exhibit complex view-dependent optical effects such as specular reflections, refraction, and low contrast. MMDOD is particularly suited for evaluating end-to-end detection-driven fusion architectures, task-driven learning strategies, and cross-sensor alignment mechanisms in challenging industrial inspection scenarios.
Database instances generated and polluted using the Perfect Pet open-source software (https://github.com/mathildemarcy/perfect_pet).
These databases contain instances of the Perfect Pet database of different sizes, polluted with various artificial unicity factors.
The clean schemas contain 13 relations: animal, animal_owner, animal_weight, appointment, appointment_service, appointment_slot, doctor, doctor_historization, microchip, microchip_code, owner, service, slot.
The polluted schemas contain 9 relations: animal, appointment, appointment_slot, doctor, microchip, microchip_code, owner, service, slot.
More information on these databases and their generation and pollution is available at https://github.com/mathildemarcy/perfect_pet.
Our current research project focuses on the cleaning for analytical use of relational databases implemented with surrogate keys and surrogate foreign keys but without natural keys. Despite their numerous benefits, surrogate keys can not only induce the presence of data quality issues within a database but also act as a major obstacle to any regular data cleaning technique, due to the artificial unicity they carry and propagate. We developed RED2Hunt, a framework dedicated to cleaning such databases, described in https://arxiv.org/abs/2503.20593.
Because of their private nature, none of the operational databases our team members worked on could be made available to the research/academic community. Thus, we decided to generate Perfect Pet, a synthetic relational database 1) to be used to facilitate the diffusion of our work on artificial unicity (the phenomenon commonly found in operational databases which resolution motivated this research project), testing and demonstrating RED2Hunt and enable the reproductibility of our experiments and results, and make it available to the community for their own use.
For these purposes, the database had to satisfy the following requirements:
Perfect Pet data should suffer from the following data quality issues:
The database was designed as a relational database supporting the appointment management application of the fictitious veterinary clinic Perfect Pet. It includes information describing the pets visiting the clinic, their owners (the clinic’s clients), the medical appointments, and the doctors working at the clinic.
Although team members never worked on a data research project related to the animal health and welfare sectors, this topic was selected to generate the synthetic data for two reasons: 1) guarantee the anonymity of the original operational databases by preventing a possible connection with the synthetic data, 2) simplify the generation process by leveraging the domain knowledge and access to an operational database in the sector of animal welfare, from one of the team members.
We thank The Jordanian Society for Animal Protection (JSAP) for allowing us to use a part of their operational database as a starting point to generate our synthetic Perfect Pet database, although it does not suffer from any of the data quality issues mentioned above.