Giant makes it easier for journalists to search, analyse, categorise and share unstructured data. It takes many file formats, indexes them (including converting images to text using OCR) and provides a UI for search. Users can upload their own files but it also scales up to terabytes of data.
Giant is part of the Guardian's "Platform for Investigations" suite, you will see references
to pfi in the code. Under development since 2017, it's written in Scala and Typescript and
is maintained by the Investigations & Reporting team.
If Giant doesn't fit your needs, check out Aleph from the OCCRP and Datashare from the ICIJ.
Giant has the following pre-requisites for local development:
Giant uses three databases, run locally in Docker through docker-compose.yaml. For local running it also uses garage as an object storage:
There are various optional dependencies needed to support extraction of different file types - you can see what these are in the Brewfile or setup.sh script
Elasticsearch requires Docker to have at least 4GB of memory from the preferences menu otherwise it will exit with no log output and error 137.
For Guardian developers:
- Janus credentials are not required to run Giant locally.
- The Giant Runbook
Install nodejs, scala, jvm - here we use mise:
mise install
Then run the setup script:
./scripts/setup.sh
Run the Scala backend:
./scripts/start-backend.sh
This will also automatically launch the databases in the background by running
docker-compose up -d.
In a separate terminal, run the Create React App frontend:
./scripts/start-frontend.sh
The frontend script will wait for the backend to start before launching Giant at
http://localhost:3000.
Once Giant has started, follow the admin quickstart guide.
We have started using dev containers to isolate dev environments from the host machine.
Giant makes use of https://github.com/guardian/devenv to simplify the dev container configuration - this will be installed
by mise if you have that, otherwise you'll need to install it manually. There's some documentation on using dev containers
in the .devcontainer README file, in the guardian/devenv repo and here https://containers.dev/.
To use the checked in devcontainer version to run giant, open Giant up in your ide of ch
To run giant inside a dev container, first it's worth generating your local devenv configuration, in case you have customised it at all:
devenv generate
Next, open up .devcontainer/user/devcontainer.json in either VS Code or IntelliJ (Note: dev container support in IntelliJ improve significantly in summer 2026, so make sure you have the latest version.)
In IntelliJ, right click on the user config file, go to 'dev containers'and then either 'Create dev container and mount sources' or 'Create dev container and clone sources'. In VScode, do shift+cmd+p and then 'reopen in dev container.
In general, the best practice is to 'create dev container and clone sources' so that the dev container is isolated from your machine as much as possible. However, in instances where you are jumping back and forth a lot between your local machine and the dev container, you may prefer to mount sources instead, so changes are synced instantly rather than having to go via github.
A new IDE window will eventually open. You'll then need to setup giant within the dev container (should be a one off). You can either use the IDE terminal for this or, get the name of the container from your Docker desktop or your IDE and run:
docker exec -it -u vscode container_name bash -l
to open a shell inside the container. The giant project will be at /IdeaProjects/giant. Note that the user is always called
vscode even if you are using IntelliJ to run the container.
Once you have a terminal in the dev container, mise install should have already happened so you can run setup.sh straight away:
./scripts/setup.sh
and then use the start-backend/frontend scripts as described above. Note that in devenv.yaml we can add port forwarding
for the bits of giant we need to access on the local machine. If you change these port numbers you'll need run
devenv generate and then restart your container - best way to do this is to close the devcontainer IDE window and then
reopen it using the local machine IDE.
You can use dev-nginx to more easily access Giant and the backing databases whilst running locally.
dev-nginx setup-app util/nginx-mapping.yml
- Giant: https://pfi.local.dev-gutools.co.uk/
- neo4j: https://neo4j.pfi.local.dev-gutools.co.uk/
- Enter
bobas the password when prompted
- Enter
- Elasticsearch: https://elasticsearch.pfi.local.dev-gutools.co.uk/
- Cerebro (to manage Elasticsearch): https://cerebro.pfi.local.dev-gutools.co.uk/
- Garage: https://garage.pfi.local.dev-gutools.co.uk/
- Username:
garage-user - Password:
reallyverysecret
- Username:
To run all unit tests:
sbt test
To run all integration tests:
sbt int:test
To run a specific integration test:
sbt 'int:testOnly controllers.api.WorkspacesITest'
Run npm test --prefix e2e-tests to test genesis account creation against disposable
databases and a separate Giant instance. Cucumber scenarios and the infrastructure
live in e2e-tests/. See the E2E setup guide
for prerequisites, ports and debugging. The GitHub Actions E2E workflow runs the
same script.
To terminate the databases without losing data:
docker-compose down
To terminate and delete data:
docker-compose down -v
The Guardian welcomes contributions to Giant. We do not yet have a publicly accessible CI server but please ensure all tests pass by running the build script locally:
./scripts/teamcity.sh
We do not yet publish deployment templates for Giant in either cloud hosts or locally. If you are interested in deploying Giant please get in touch by raising a GitHub issue on this repository.
https://docs.google.com/drawings/d/1wcTY9KLhkYqxmwzsyZ3DsWcc0v-ax5kMKWtYb4HZgF0
Giant uses the Apache 2.0 licence. Some libraries used are licensed separately:
.rararchives (v4 and below).ziparchives.emlRFC 5322 emails.mboxemail archives.msgOutlook email files.pstOutlook email archives.olmOutlook for Mac email archives/backups.png,.jpg,.tiffimages (including OCR).pdf(including OCR)- Microsoft Office Word, Excel and Powerpoint files
- Various plain text files (see DocumentBodyExtractor)
- Audio files
- fully supported
.wav.mpeg.opus.caf.mp4.aac(tika sometimes has trouble detecting these)
- transcribed but preview doesn't work
.aff.amr.wma
- fully supported
- Video files
- fully supported
.mov,.qt.m4v.3gpp.mp4
- transcribed but preview doesn't work
.flv.wmv.msvideo.mpeg
- fully supported
Experimental features are enabled through feature flags in the Settings page:
- New UI: a simplified UI implemented using the Elastic UI toolkit
- Page Viewer: a unified document viewer showing text, OCR and search highlights inline on the original document
In addition to any contributors named in this repository, the following contributed to Giant whilst it was closed source at the Guardian:

