METHODOLOGY
Methodology
The more than 100,000 leaked documents reviewed for this research were obtained from a
source who also shared the data with the following partners: Amnesty International, Justice
For
Myanmar,
Paper
Trail
Media,
The
Globe
and
Mail,
the
Tor
Project,
the
Austrian
newspaper DER STANDARD and Follow The Money. These partners, as well as the authors
of this report, collaborated to analyze, verify, and report on various aspects of this data.
The InterSecLab team developed and hosted a collaborative infrastructure to facilitate a
multi-partner, cross-border investigation of the collection of documents included in the
data leak. We deployed an OpenSearch cluster, indexing the documents into the opensource document discovery platform Datashare. Our consortium of partners used this
platform to query the documents. Partners used this infrastructure to share additional data
and findings from interviews with human sources and other subscription-based databases
including corporate records, patents, etc.
We then used the open-source text extraction library Apache Tika as well as a combination
of Tesseract and Apple Vision framework for optical character recognition and ran every
document through the Llama family of large language family of models from Meta to
translate and generate an English language summary of each document. Our research team
then indexed those summaries into a modified version of the open source wiki.js software
which allowed us to select documents.
Selected documents were further translated and analyzed. Some of the data was already
available in English (as it is often the case that Geedge Networks communicates with its
international clients in English). When documents were only available in Chinese, we crossreferenced
multiple
machine
translations
and
large
language
models.
InterSecLab
consulted with Chinese speakers and experts on the interpretation of key documents used
in this research. Additional analysis was made by reviewing Geedge’s illustrations and
network diagrams. Details from our findings were cross-referenced using import export
databases, LinkedIn, and public news sources, to corroborate the events described in the
leaked documents. The leak contains source code, and we have indexed the source code.
It is important to note that for this research InterSecLab did not conduct a comprehensive
review of the source code.
11