For further refinement: In retrospect, the short time frame for recruiting and training researchers led to some inconsistencies in the research results. In particular, the legal adviser’s review revealed that not all researchers demonstrated the same understanding of the level of detail being requested, which has led SMEX to conduct additional rounds of review. In future, we recommend that, when resources are available, in-person trainings on the research methodology and workbook should be organised and attendance should be a condition of payment. A longer, multi-round recruitment process, with some kind of assessment to measure the researcher’s capacity and eye for detail, would also be useful and help expedite data review and verification. Data collection and review: Findings and challenges After five months’ preparation, data collection began in early March 2017. Researchers were given one month to complete the original research process and one month to complete their peer review, which involved checking the folder and workbook of a second country. Three researchers dropped out before the research was complete for health and family reasons. Meanwhile, one researcher revealed late in the process that they did not read Arabic. Also, because some researchers were behind schedule, the peer review process was also delayed. Ultimately, the first round of original research and peer review concluded in June 2017. In July 2017, SMEX and the legal adviser conducted an overall review of all the workbooks. In all, the law catalogues grew from 142 in the first dataset to around 240, the vast majority of them with official or unofficial translations. Dozens of key provisions were identified. Several draft laws were noted, and case law, a completely new type of information in this version of the ADRD, was identified in six countries.52 Following a final review by SMEX and in-country experts, the expanded datasets will be made public. For further refinement: As mentioned above, SMEX has added two more rounds of review to ensure that the data we have is as accurate and upto-date as possible. Unfortunately, this has delayed making the data available, which could also compromise its accuracy, if too much time passes. To avoid such delays in the future, we recommend that 52 Case law was identified in only six countries: Egypt, Jordan, Kuwait, Lebanon, Morocco and Mauritania. research supervisors implement a phased approach with interim milestones. For example, data could be collected, reviewed and verified for one worksheet at a time and combined with periodic group calls to raise and resolve concerns or challenges encountered. This would not only help ensure that researchers develop a shared understanding of the nuances of the research process but will also yield better results that can be publicised more quickly. Finally, while we included draft laws and provisions and case law in the current workbook in response to stakeholder requests for this data, we are delaying their integration into the public dataset pending more detailed research and review. Gathering data about case law posed several problems with regard to not only locating and sourcing decisions but also in developing a consistent approach to explaining how cases interpret the relevant laws, which is essential to being able to publish authoritatively on their impact. In subsequent phases of the project, we will explore addressing such challenges by integrating into the methodology existing approaches to analysing case law, such as that of Columbia University’s Global Free Expression Case Database.53 The future roadmap Perhaps unlike other research methodologies, the one for the Arab Digital Rights Datasets was also designed to be expressed as a data model, or a conceptual framework to organise and standardise the data collected. Rendering the methodology as a data model makes it much easier to share, extend, combine and repurpose information, especially by machines. In parallel with the data collection and review process, we worked with technologist Seamus Tuohy to create the data model for the ADRD and a related API, or application programming interface. An API is a piece of code that sits between a database and a graphic user interface (GUI) that calls information from the database according to what a user needs. This data model and API will be used to build a database of the Arab laws collected and make the data both human and machine-readable. But it is our hope that these technical interpretations of the methodology will also afford other organisations conducting similar research the opportunity to make their data more available and accessible too. To this end, SMEX is now forming a working group to explore the potential for this data model to become a global standard for aggregating, organising 53 https://globalfreedomofexpression.columbia.edu/cases 16 / Unshackling Expression

Select target paragraph3