We create data from humanities sources for several reasons. Humanities sources are historical products of human design and interaction, and as such, data are vibrant with possibility (Hoekstra & Koolen, 2019; Padilla, 2016). Padilla (2016) further argues that humanities scholars create data from humanities sources because humanities scholars are builders; we build infrastructure and collections, and we make visible less visible structures represented in digital humanities project collections. Creating data from humanities sources requires the interplay of data mined from historical documents and materials and advanced research and digital humanities tools, enabling digital scholars to capture, interpret, and understand facets of human interaction (Posner, 2015).
The types of questions that humanities data can answer that other methods cannot include identifying trends in large datasets. Humanities data can also answer questions about trends in the data, including the identification of networks of influence on the lives of specific groups. These networks of influence are evident in the data projects reviewed, including the Tudor Networks, the Early African American Films, 1909-1930, and the Colonial Probate Records in the Fairfax Court Slavery Index projects (Ahnert & Ahnert, 2023; Bollinger et al., 2023; Cifor et al., 2018; Hoekstra & Koolen, 2019). Digital humanities work can also reveal implicit biases in archival material and how these biases are constructed and perpetuated in scholarship (Guldi, 2024; Padilla, 2016).
During our assigned data creation activity, which used the Alabama Slave Narratives from the Federal Writers’ Project (1936-1938), I encountered challenges working with humanities sources that were not consistently complete. I also recognized the importance of documenting my decision-making process to ensure transparency and community access.
The concept of transparency is essential in the collection-building process. It includes documenting and critiquing how researchers process data, determine which data to include and exclude, assess the representativeness of the data, and identify the organizational biases reflected in the data (Padilla, 2016). A data scope guides this process, which Hoekstra and Koolen (2019) define as “a coherent set of methodological principles that characterize the interaction between researchers and their data and the transformation of a cluster of data into a research instrument” p.80. This ‘methodological argumentation’ of data scopes is not only the process of documenting the steps involved in transforming data, the rationale for those steps, and their consequences, but also a critical examination of the datasets (Hoekstra & Kollen, 2019). Researchers should view data scopes as interpretive, iterative processes and include the following steps: selection, modelling, normalization, linking, and classification (Hoekstra & Kollen, 2019).