Phishy Mailbox is a tool for researching human factors of phishing. It was created for researchers to easily run phishing studies using real emails in an in-basket exercise, where participants categorize the emails into a number of configurable folders.
Participant interface, showcasing a user-friendly design to categorize emails.In depth documentation is available in both english and german.
- Install Docker on your machine. For Windows, Docker Desktop is recommended. Keep in mind you need admin privileges to execute Docker.
- Download the docker-compose.yml from this repository and place it into an empty folder.
- Download the docker image from Dockerhub into the same folder.
- Start docker, if necessary.
- Start the command line interface and navigate to the folder containing image and yml file.
- type in: docker compose up -d
- Wait for the program to load
The application should start and be reachable from localhost:3000 (user interface) or localhost:3000/admin.
If you use this software for your research, please don't forget to cite it in your papers! Link to the publication: https://www.ndss-symposium.org/wp-content/uploads/usec25-37.pdf
The first version of this tool was created in the context of a bachelor's thesis at the department for usable security and privacy at Leibniz Universität Hannover.
We welcome contributions from the community. Feel free to open issues and submit pull requests.
Prerequisites: Docker and Yarn
The application consists of two components. The first one is a PostgreSQL database that can be launched after installing docker via running docker compose -f docker-compose.dev.yml up -d in the root directory.
Afterwards you can run the following commands to start the Next.js server that serves both the spa-frontend as well as the backend API using prisma as the ORM.
yarn
yarn prisma generate
yarn prisma db push
yarn node ./prisma/seed.mjs
yarn devThe same two components used for development are also required for deployment, general instructions to deploy a next.js application are available here. During development a deployment using Vercel and supabase was tested and can be recommended.
When upgrading an existing deployment across a PostgreSQL major version (e.g. the move from 15 to 17), the bundled database needs a dump & restore — see UPGRADING.md.
The test suite is split into two layers, both running against a dedicated test
database (PostgreSQL on port 5434, separate from the dev DB on 5432). Start it
once with:
yarn test:db:up # docker-compose.test.yml, Postgres on :5434
# ... run tests ...
yarn test:db:down # stop and remove itThese call the tRPC routers directly (appRouter.createCaller) against the test DB.
Every test runs inside its own transaction that is rolled back afterwards, so
nothing is ever committed — tests are completely isolated and parallel-safe (each
Vitest worker uses its own connection). This is the reliable coverage metric to
iterate on.
yarn test:unit # run once
yarn test:unit:watch # watch mode
yarn test:coverage # with coverage -> coverage/unit/index.htmlTests live next to the code as src/**/*.test.ts; shared helpers/factories are in
test/integration/.
Browser tests of the real user flows. For complete isolation with parallelism, each
Playwright worker boots its own Next.js server on its own port and clones its own
database from a migrated template; before every test the worker DB is truncated and
re-seeded (clean slate). The first run performs a next build.
yarn test:e2e # build + run (2 workers by default)
PLAYWRIGHT_WORKERS=4 yarn test:e2e # more parallelism (one server/DB per worker)
SKIP_BUILD=1 yarn test:e2e # reuse an existing .next buildSpecs and fixtures live in test/e2e/.
Coverage is measured at the integration layer (yarn test:coverage), which maps
cleanly and reliably to src/server/**. The way to raise coverage is to add more
integration tests there. The e2e suite intentionally does not produce a coverage
number: it runs against the compiled Next.js production server, whose V8 coverage does
not map back to source without server source maps, so any figure would be misleading.
E2E tests exist to verify the real user flows end to end, not to move a coverage gauge.

