Skip to content

Latest commit

聽

History

History
53 lines (35 loc) 路 4.41 KB

File metadata and controls

53 lines (35 loc) 路 4.41 KB

Datasets

As mentioned in README.md, this directory has to be set up to store the datasets used by mdif.

There are two directories and the following sections explain how to set them up!

Note

Most of the datasets are split by 80% for training and validation and 20% for testing! The training and validation dataset is then split again by 80% for training and 20% for validation in the training/* modules


raw

The raw directory holds the datasets we manually download and extract. The data in this directory is used by the preprocessing/compute_features module for preprocessing.

Setup

Download and extract the datasets from the following sources:

Dataset Source Remarks
CIFAKE Kaggle This dataset only contains the REAL/* (real) and FAKE/* (fake)
Unbiased Tiny GenImage Kaggle We only use the Nature/* (real), Midjourney/* (fake), and glide/* (fake) for training and testing
AutoSplice Github We only use the Authentic/* (real) and Forged_JPEG90/* (inpainted) for training and testing
CocoGlide Github This dataset contains real/* (real) and fake/* (inpainted) for training and testing
SAGI Kaggle This dataset contains original/* (real), brushnet/* (inpainted), controlnet/* (inpainted), and removeanything/* (inpainted) for training and testing

Note

The CIFAKE and SAGI dataset, by default, comes split as train/* and test/* datasets but the others do not. To address this, the preprocessing/compute_features module splits the datasets into training and testing datasets while preprocessing!

The SAGI Dataset comes with two folders, /coco/* and /raise/*. The model has only been trained on the /raise/* split of the dataset.

Once the datasets have been stored in the raw directory, head on over to the mdif notebook to see how to perform the preprocessing steps!


processed

The processed directory is further split into a train and test directories.

The files in these directories are resized image (.jpg) files as well as their corresponding extracted feature (.npy) files by the preprocessing/compute_features module.

These files have a consistent naming scheme <label>_<original_file_name>.* where the labels are as follows:

Label Description Dataset
0 Real CIFAKE REAL, AutoSplice Authentic, CocoGlide real, GenImage Nature, SAGI original
1 Fake CIFAKE FAKE, GenImage Midjourney/glide
2 Forged/Inpainted AutoSplice Forged_JPEG90, CocoGlide fake, SAGI brushnet/controlnet/removeanything