This document provides a comprehensive description of all three MongoDB collections used in the JudgeGPT research platform: participants, results, and fragments. Each attribute is explained with its data type, valid range, and meaning.
Database Name: realorfake
Collections:
- participants - Stores demographic and psychographic data about study participants
- results - Stores participant responses/judgments for each news fragment
- fragments - Stores news content (both human-written and AI-generated)
Stores comprehensive demographic and psychographic information about each study participant collected during the initial survey form.
| Attribute | Data Type | Range/Values | Description |
|---|---|---|---|
_id |
ObjectId | MongoDB auto-generated | Unique document identifier (MongoDB internal) |
ParticipantID |
String (UUID hex) | 32-character hexadecimal | Unique identifier for each participant (generated via uuid.uuid4().hex) |
ISOLanguage |
String | "en", "fr", "de", "es" |
ISO language code for participant's preferred language: • "en" = English• "fr" = French• "de" = German• "es" = Spanish |
Age |
Float | Configurable (default: 16.0 - 133.0) | Participant's age in years. Range can be customized via URL query parameters (?min_age=X&max_age=Y) |
Gender |
String | "Male", "Female", "Other", "Prefer not to say" |
Participant's self-identified gender |
PoliticalView |
Float | 0.0 - 1.0 (7-point scale) | Participant's self-assessed political orientation on a normalized 0-1 scale: • 0.0 = Far Left• 0.2 = Left• 0.4 = Center-Left• 0.6 = Center-Right• 0.8 = Right• 1.0 = Far Right• 0.5 is used as placeholder (not valid for submission) |
IsNativeSpeaker |
String | "Yes", "No" |
Whether the participant is a native speaker of the selected language |
EducationLevel |
String | "None", "High School", "Apprenticeship", "Bachelor's Degree", "Master's Degree", "Doctoral Degree" |
Highest level of education completed by the participant. Mandatory field. "None" means no formal education, not a missing value: it is the first of six options in app.py (German locale "Keine"). Parse with keep_default_na=False, otherwise pandas silently converts it to NaN and undercounts the category. |
NewspaperSubscription |
Float | 0.0, 1.0, 2.0, 3.0 | Number of newspaper subscriptions the participant has: • 0.0 = None• 1.0 = One• 2.0 = Two• 3.0 = Three or more• 1.5 is used as placeholder (not valid for submission) |
FNewsExperience |
Float | 0.0 - 1.0 (7-point scale) | Participant's self-assessed experience/familiarity with fake news: • 0.0 = Completely unfamiliar• 0.2 = Mostly unfamiliar• 0.4 = Somewhat unfamiliar• 0.6 = Somewhat familiar• 0.8 = Mostly familiar• 1.0 = Completely familiar• 0.5 is used as placeholder (not valid for submission) |
IpCountry |
String | Country name or "" |
Country derived from the geolocation lookup. Empty when the lookup failed. |
IpContinent |
String | Continent code or "" |
Continent derived from the same lookup. Empty when the lookup failed. |
RecruitmentRoute |
String | "direct", "age-restricted", "challenge-link", "campaign", "language-preset", "other" |
The kind of link a participant arrived through, derived from the URL query parameters. The parameters themselves are not stored. |
Earlier versions of the application stored four further fields: ScreenResolution, IpLocation (the complete geolocation response, including the address itself, coordinates to six decimals, postal code, city, region and timezone), UserAgent, and QueryParams (every URL parameter verbatim).
None of them were used by any analysis, and in combination they identified individuals: across the collected sample, 208 of 249 distinct (address, user agent, screen resolution) triples occurred exactly once, and the stored query parameters carried the age bounds that mark the secondary-school cohort as a group.
app.py now reduces the geolocation response to IpCountry and IpContinent and the query parameters to RecruitmentRoute before the database write, and no longer requests the user agent or screen resolution at all. The raw values exist only for the duration of the request and are never persisted. This is data minimisation at collection rather than a filter applied when publishing, so the database, the published deposit and this document describe the same set of fields.
Records collected before this change may still carry the four original fields. They are withheld from the published deposit in every case.
Stores individual participant responses/judgments for each news fragment they evaluate. This is the core collection containing all judgment data.
| Attribute | Data Type | Range/Values | Description |
|---|---|---|---|
_id |
ObjectId | MongoDB auto-generated | Unique document identifier (MongoDB internal) |
ResultID |
String (UUID hex) | 32-character hexadecimal | Unique identifier for each individual response/judgment (generated via uuid.uuid4().hex) |
ParticipantID |
String (UUID hex) | 32-character hexadecimal | Foreign key linking to the participant who made this judgment (matches ParticipantID in participants collection) |
FragmentID |
String (UUID hex) | 32-character hexadecimal | Foreign key linking to the news fragment being evaluated (matches FragmentID in fragments collection) |
HumanMachineScore |
Float | 0.0 - 1.0 (7-point scale) | Participant's judgment on whether content is human or machine-generated: • 0.0 = Definitely Human Generated• 0.2 = Probably Human Generated• 0.4 = Likely Human Generated• 0.6 = Likely Machine Generated• 0.8 = Probably Machine Generated• 1.0 = Definitely Machine GeneratedLower values = more human-like perception |
LegitFakeScore |
Float | 0.0 - 1.0 (7-point scale) | Participant's judgment on content legitimacy: • 0.0 = Definitely Legit News• 0.2 = Probably Legit News• 0.4 = Likely Legit News• 0.6 = Likely Fake News• 0.8 = Probably Fake News• 1.0 = Definitely Fake NewsLower values = more legitimate perception |
TopicKnowledgeScore |
Float | 0.0 - 1.0 (7-point scale) | Participant's self-assessed familiarity with the topic: • 0.0 = Not at all familiar• 0.2 = Slightly familiar• 0.4 = Somewhat familiar• 0.6 = Fairly well familiar• 0.8 = Very well familiar• 1.0 = Extremely well familiar |
Timestamp |
String (ISO 8601) | ISO format datetime string | Exact timestamp when the response was submitted (format: "YYYY-MM-DDTHH:MM:SS.mmmmmm"), generated via datetime.now().isoformat() |
TimeToAnswer |
Float | Positive number (seconds) | Time elapsed (in seconds) between when the fragment was displayed and when the participant submitted their response. Calculated as (end_time - start_time).total_seconds() |
SessionCount |
Integer | Positive integer (1, 2, 3, ...) | Sequential response number within the participant's current session (starts at 1, increments with each submitted response) |
Origin |
String | "Human", "Machine" |
Ground truth: Actual origin of the news fragment: • "Human" = Written by a human journalist• "Machine" = Generated by AI/LLM |
IsFake |
Boolean | true, false |
Ground truth: Whether the content is actually fake news/misinformation: • true = Content is fake/misinformation• false = Content is legitimate news |
ReportedAsBroken |
Boolean | true, false |
Whether the participant flagged this fragment as technically incorrect or broken: • true = Participant reported an issue• false = No issues reportedFragments reported as broken bypass the validation requirements for scores |
For Analysis:
- HumanMachineScore >= 0.5 suggests participant perceived content as machine-generated
- HumanMachineScore < 0.5 suggests participant perceived content as human-generated
- LegitFakeScore >= 0.5 suggests participant perceived content as fake
- LegitFakeScore < 0.5 suggests participant perceived content as legitimate
Accuracy Calculations (as implemented in aggregate_results()):
- HM_Accuracy:
(HumanMachineScore >= 0.5) == (Origin == "Machine") - LF_Accuracy:
(LegitFakeScore >= 0.5) == IsFake
Stores news content fragments used in the study, including both human-written articles and AI-generated content.
| Attribute | Data Type | Range/Values | Description |
|---|---|---|---|
_id |
ObjectId | MongoDB auto-generated | Unique document identifier (MongoDB internal) |
FragmentID |
String (UUID hex) | 32-character hexadecimal | Unique identifier for each news fragment (generated via uuid.uuid4().hex) |
Content |
String | Text content | The actual news fragment text that participants evaluate. Can be: • Human-written news excerpt from real outlets • AI-generated content from LLMs Typically 1-3 paragraphs in length |
Origin |
String | "Human", "Machine" |
Source type of the content: • "Human" = Written by human journalist• "Machine" = Generated by AI/LLM |
ISOLanguage |
String | "en", "fr", "de", "es" |
ISO language code of the content: • "en" = English• "fr" = French• "de" = German• "es" = Spanish |
IsFake |
Boolean | true, false |
Whether the content constitutes fake news/misinformation: • true = Fake news/misinformation• false = Legitimate news |
CreationDate |
Date/DateTime | ISO datetime | Date when the fragment was created/ingested into the database |
HumanOutlet |
String | Outlet name or empty string | Only for Origin="Human": Name of the news outlet/publication source (e.g., "New York Times", "BBC") • Empty string "" if Origin="Machine" |
HumanURL |
String | URL or empty string | Only for Origin="Human": Original URL of the source article • Empty string "" if Origin="Machine" |
MachineModel |
String | Model identifier or empty string | Only for Origin="Machine": Identifier of the LLM used to generate the content. Examples: • "openai_gpt-4"• "openai_gpt-35-turbo"• "openai_gpt-4o_2024-08-06"• "meta_llama-2-13b-chat"• "google_gemma-7b-it"• "mistralai_mistral-7b-instruct-v0.2"• "microsoft_Phi-3-mini-4k-instruct"• Empty string "" if Origin="Human" |
MachinePrompt |
String | Prompt text or empty string | Only for Origin="Machine": The exact prompt used to generate the content (e.g., "Write a tweet about '''climate change''' in English in the style of CNN") • Empty string "" if Origin="Human" |
| [Dynamic Keys] | Various | Varies by fragment | Additional metadata fields from RogueGPT's prompt template components (e.g., "Format", "SeedPhrase", "Language", "Style") are stored dynamically based on the generation configuration |
When retrieving fragments for a participant, the system uses:
{
"$match": {"ISOLanguage": <participant_language>},
"$project": {"FragmentID": 1, "Content": 1, "Origin": 1, "IsFake": 1, "_id": 0},
"$sample": {"size": 50}
}This randomly selects 50 fragments in the participant's chosen language.
-
Participant Registration:
- User fills out demographic form
- Record inserted into
participantscollection ParticipantIDgenerated and stored in session
-
Fragment Retrieval:
- System queries
fragmentscollection based onISOLanguage - 50 random fragments loaded for the session
- System queries
-
Judgment Collection:
- For each fragment, participant provides three scores
TimeToAnswercalculated from presentation to submission- Record inserted into
resultscollection with:- Link to
ParticipantID - Link to
FragmentID - All judgment scores
- Ground truth (
OriginandIsFake)
- Link to
-
Feedback Loop:
- Every 5 responses, system calculates accuracy metrics
- Aggregation performed on session's
resultsdata - No database writes for feedback calculations
- All demographic fields must be selected (cannot remain at placeholder values)
PoliticalView,NewspaperSubscription, andFNewsExperiencecannot be0.5or1.5(placeholder values)- Consent toggle must be
true Agemust be within configured min/max range
- If
ReportedAsBroken = false, all three scores must be selected (cannot remain at0.5) - If
ReportedAsBroken = true, validation is bypassed TimeToAnswermust be positiveParticipantIDandFragmentIDmust reference existing documents
Contentcannot be emptyOriginmust be either"Human"or"Machine"- If
Origin = "Human":HumanOutletandHumanURLshould be populated,MachineModelandMachinePromptshould be empty - If
Origin = "Machine":MachineModelandMachinePromptshould be populated,HumanOutletandHumanURLshould be empty
The corpus was checked against these rules for the v1.2.0 deposit. One machine-origin fragment carried a bare domain in HumanURL with an empty outlet, a data-entry artifact; the field was cleared and the row kept. No other violation exists in the released data.
To analyze the full dataset, merge the three collections:
# Merge results with fragments
df_merged = df_results.merge(df_fragments, on="FragmentID", suffixes=("_result", "_fragment"))
# Merge with participants
df_merged = df_merged.merge(df_participants, on="ParticipantID", suffixes=("", "_participant"))- Demographic Predictors:
Age,Gender,PoliticalView,EducationLevel,FNewsExperience - Judgment Accuracy: Compare participant scores to ground truth (
Origin,IsFake) - Model Performance: Group by
MachineModelto compare which LLMs fool participants most effectively - Temporal Effects: Use
SessionCountandTimeToAnswerto detect fatigue or learning effects - Cross-Cultural: Compare judgment patterns across
ISOLanguagegroups
- ParticipantID is a randomly generated UUID with no personally identifiable information
- IP geolocation is approximate (city/country level) and collected via third-party API
- User agent and screen resolution are technical metadata, not personal identifiers
- All data collection follows informed consent procedures (consent toggle required)
- Participants can flag problematic fragments via
ReportedAsBroken
- Platform: JudgeGPT v1.0.3
- Generator: RogueGPT v0.9.3
- Database: MongoDB (realorfake)
- Last Updated: October 2025