Category Archives: Uncategorized

Text Analysis with Voyant

Exploring the assigned case studies, America’s Public Bible, Signs@40, and Robots Reading Vogue, in conjunction with the analyses by Goldstone (2014), Leonard (2014), and Mullen (2021), highlights the effectiveness of Voyant tools for creating dynamic text mining experiences for text analysis projects. These case studies demonstrate how Voyant can produce dynamic, interactive data visualizations for users, as in the Signs@40 project, which features a network graph representing author cocitations and an interactive, curated table of contents with commentaries from past, current, and future editors and contributors.

By applying the text analysis algorithms built into Voyant tools, I analyzed the entire corpus or specific state files of the WPA Slave Narratives project. The content tools allowed me to view word use by state and compare frequently used words between two states. I noticed that most interviews across states used some of the same words at high frequency, including “old”, “come”, and “got”, but within states, these words varied by region.

Voyant tools also enabled me to understand narrators’ words in context across the corpus. For example, the word “got” was used in interviews across the corpus in contexts referring to actions done to them as persons with limited agency.

The Voyant tool also identified distinctive words that varied in frequency due to dialect differences, topic relevance, or the interview format. By connecting frequently used words across state datasets, I gained insights into the meanings of these differences. Further, text analysis of these interviews using Voyant tools revealed a deeper understanding of how variations in word use reflected narrators’ experiences as formerly enslaved individuals.

A Voyant feature most helpful for text analysis is its ability to remove common words, or stop words, that can clutter the analysis. I was able to customize the stop word list by adding corpus stop words to Voyant’s default list, enabling me to refine my visualizations. For example, one of the most used words by narrators in Georgia appeared to be “war” as one of the top five most frequently used words generated in the Cirrus tool from the corpus was “war”. When I examined the context in which the word was used, however, it was clear that the actual word counted was the commonly used word “warn’t,” so I added “warn’t” to the stop word list to refine the distinctive words list results.

The assigned WPA Slave Narratives project answered the question of how narrators’ words used to describe their enslavement experiences and spoke to the nature of the institution of slavery in some American states over time. This project also helped answer the question of how narrators’ accounts of enslavement experiences were used and changed over time, similar to findings from the America’s Public Bible project, which explored how the Bible, as a text, was used in the public sphere and evolved over time (Guldi, 2024). As a researcher, I find this text-mining tool invaluable for analyzing corpora and making meaningful comparisons across datasets.

Humanities Data

We create data from humanities sources for several reasons. Humanities sources are historical products of human design and interaction, and as such, data are vibrant with possibility (Hoekstra & Koolen, 2019; Padilla, 2016). Padilla (2016) further argues that humanities scholars create data from humanities sources because humanities scholars are builders; we build infrastructure and collections, and we make visible less visible structures represented in digital humanities project collections. Creating data from humanities sources requires the interplay of data mined from historical documents and materials and advanced research and digital humanities tools, enabling digital scholars to capture, interpret, and understand facets of human interaction (Posner, 2015).

The types of questions that humanities data can answer that other methods cannot include identifying trends in large datasets. Humanities data can also answer questions about trends in the data, including the identification of networks of influence on the lives of specific groups. These networks of influence are evident in the data projects reviewed, including the Tudor Networks, the Early African American Films, 1909-1930, and the Colonial Probate Records in the Fairfax Court Slavery Index projects (Ahnert & Ahnert, 2023; Bollinger et al., 2023; Cifor et al., 2018; Hoekstra & Koolen, 2019). Digital humanities work can also reveal implicit biases in archival material and how these biases are constructed and perpetuated in scholarship (Guldi, 2024; Padilla, 2016).

During our assigned data creation activity, which used the Alabama Slave Narratives from the Federal Writers’ Project (1936-1938), I encountered challenges working with humanities sources that were not consistently complete. I also recognized the importance of documenting my decision-making process to ensure transparency and community access.

The concept of transparency is essential in the collection-building process. It includes documenting and critiquing how researchers process data, determine which data to include and exclude, assess the representativeness of the data, and identify the organizational biases reflected in the data (Padilla, 2016). A data scope guides this process, which Hoekstra and Koolen (2019) define as “a coherent set of methodological principles that characterize the interaction between researchers and their data and the transformation of a cluster of data into a research instrument” p.80. This ‘methodological argumentation’ of data scopes is not only the process of documenting the steps involved in transforming data, the rationale for those steps, and their consequences, but also a critical examination of the datasets (Hoekstra & Kollen, 2019). Researchers should view data scopes as interpretive, iterative processes and include the following steps: selection, modelling, normalization,  linking, and classification (Hoekstra & Kollen, 2019).

Why Metadata Matters

Metadata is critical in digital humanities because, at its core, it is descriptive and contextual information about a digitized item that enables discovery (Martin & Neatrour, 2015). The field of digital humanities relies on the perpetuation of knowledge from the past; therefore, metadata serves as the lifeline for this process. Further, metadata is an essential means of organizing collections and connecting items across collections.

The rise of AI integration poses a threat to library infrastructure and discovery systems. The push for AI use in library systems, combined with budgetary constraints and the consolidation of library catalog vendors, threatens the accuracy and diversity of metadata (Olsen, 2025). Additionally, the removal of human labor experts threatens the accuracy and diversity of metadata, as humans make choices about what information to record.

Based on my experience with the “Digitizing Your Kitchen” assignment, describing items in Tropy and displaying them in Omeka was labor-intensive. I spent significant time ensuring that images were handled and labelled in accordance with the Library of Congress and Dublin Core standards, and that the metadata I entered into Tropy is publicly searchable. The templates and vocabularies developed by Tropy and Omeka, however, facilitated this process in terms of accuracy and time. Tropy allowed me to add detailed metadata to my digitized images and to group them in my project. Omeka, another tool used in this project, provided the digital platform for describing and exhibiting my project images. Through a complex, labor-intensive process, I organized and optimized my digitized kitchen objects to improve their searchability for other researchers.

Metadata matters in digital humanities, and decisions about metadata creation are best made by humans. The rise of AI in the digitization process threatens human labor, as AI leaders push to aggregate and monetize standardized metadata. Ultimately, the push to use AI in this process undermines the archival value of human curation.

African American Newspapers, 1827-1998 Review

African American Newspapers, 1827-1998

Database Review

Product Overview and Description

The African American Newspapers 1827-1998 database is a digitized collection of 282 U.S.-based newspapers that cover African American history and culture during this specified period. This unique collection offers comprehensive access to press coverage of African American life during the Antebellum South, the spread of abolitionism, the emergence of the Black church, the Jim Crow Era, the Great Migration, the Harlem Renaissance, the Civil Rights Movement, and the leaders of these periods. This archive provides full, digitized access to a wide range of newspapers published in 40 states. Most newspapers in this archive are limited to one or only several issues for the 171 years of articles covered.

History. Provenance. Digitization Background

Provided by Readex (a division of NewsBank, Inc.), this database draws on over 80 years of institutional expertise. According to the Readex website, newspapers in this archive are “newly digitized” and represent the most extensive African American newspaper archives in the world. The materials are amassed (and perhaps digitized) from major archival collections, including the Wisconsin Historical Society, the Kansas State Historical Society, and the Library of Congress. 

This collection is curated based on James Philip Danky’s (1998) book: African-American Newspapers and Periodicals: A National Bibliography. This collection can be cross-searched with other Readex collections, including America’s Historical Newspapers, Black Life in America, and various era-specific collections that chronicle African American life and culture.

User Interface, Navigation, and Searching

The interface and navigation of the African American Newspapers 1827-1998 database mirror the style of other Readex Historical Newspapers, including Caribbean American, America’s Historical Newspapers, and the American Business Mercantile Newspaper archive. These standards include a Hero Image design style with a top bar and top menu.

Users can search the newspaper database through a basic or advanced search. The basic search allows users to enter keyword (s) or phrase words to generate results matching search terms within a single newspaper or across multiple newspapers. In addition to keyword searches, users can search articles by date range and geographic location. Users can also search by publication name, publication location, article type, language, eras in American history, U.S. Presidential era, decade, and a specific year. Users can browse full-text, complete issues of periodicals in this database.

Search results provide image reproductions of the original newspaper. Once you identify search results, users can preview the article portion with the search word or phrase highlighted in yellow and view the specific article in the context of its full-text publication issue. Users can download or export images as PDFs, cite sources generated by searches, and receive automatically generated, copyable citations in various formats. The user can email search items to one or more recipients. The user can save the search separately or alongside other searches by creating a personal folder account via the “My Folder” link in the top-right header.

The user can also access the Readex Text Explore tool, integrated into database searches, to visualize linguistic and thematic patterns within a single periodical or across multiple periodicals through its snapshot view of the Voyant Tool. By selecting one or more search results, users can click the “Explore” tab to the right of the “Save” button. Explore activates the Voyant Tool, which enables users to perform computational analysis of search terms, including term clustering, frequencies, and trends, and to see linguistic and thematic patterns in a single work or across multiple works. Voyant Tool is flexible and allows modification of search terms, with real-time updates to the visualization panel. Users can sort results by best match, newest, and oldest content. The articles reviewed are indexed, which helps with cross-database searches.

The top bar includes Help, Change Databases, and Share Feedback buttons. Using the top bar features, users can access the Help button. This seventeen-page document, comprising text and images, provides in-depth instructions on searching the database, working with search results, and using documents, as well as the terms of use, privacy policy, and browser requirements. Left navigators offer options for customizing the visualization of corpus documents. Users can export findings in various formats, including URLs, HTML snippets, or images. User feedback buttons appear at the top and the bottom bars of the database. The option to toggle between databases is also at the top bar. This database does not include an audio feedback option.

Access

This database partners with “renowned repositories, libraries, and historical societies to create authoritative digital collections” that are accessible by subscription. The Sales Department provides additional services, including educational lesson plans. Users must abide by the terms of use when accessing the database. All content included on this database is Readex property or the property of their content providers and protected by copyright law.

Review and Reception

The Readex website posts database reviews under the “Impact in the Classroom” and “Recent Praise” links. The website did not post reviews for this particular database. A web search unearthed a 2010 review of the Wagner College Horrmann Library, posted on its library blog, which describes a “unique online collection that offers unprecedented insights into African American history, culture, and daily life”.

Critical Evaluation

Scholars of history, African American life, and culture would likely appreciate this resource, though most of the newspapers in this database offer only partial collections. Some search results are clearer than others due to the age of the newspapers, variations in typeface and font, and the digitization and OCR conversion processes, which affect the readability of text from earlier publications. Additionally, black newspapers of ten U.S. states are missing from this collection, and some states are overrepresented, including California, the District of Columbia, Illinois, and Kansas.

The database was relatively easy to navigate. A more prominent ‘Help’ button and instructions on how to use the Text Explorer feature, which provides the desired Digital Humanities tools, would have made the database easier to use. The key to successfully accessing and using the Text Explorer tools in this database is to view the five-minute promotional video demonstration produced by Readex and posted to the Readex Blog on September 08, 2020.

Issues Observed

The Explore option is identified on the database as a “new” feature prominently displayed in red letters to the right of the “Text Explorer” tab. Instructions for using the Text Explorer are missing from the Help PDF. Readex should update the Help feature to allow users to access the Text Explorer demonstration video and remove Internet Explorer, an obsolete web browser, from the list of required browsers.

A Guide to Digitization

When an artifact is digitized, the goal is to create consistent images of documents by converting an analogue signal or code into a digital signal or code (Terras, 2012). Through this process, any items accessible through photography, sound, and moving images can be digitized using various techniques (Terras, 2012). For example, in this module, I captured photographs of items from my kitchen for digitization, including still photos and videos of a recipe, a bowl of fruit, and a container of shelled and jarred beans. The representations I captured of these items varied, with the video versions offering a more comprehensive, contextualized view than the photographs.

Additionally, while the goal of digitization is to create consistent images of documents and artifacts, the process relies on artifacts that are “fundamentally individual and inconsistent” (Terras, 2012). Therefore, I cannot capture every aspect of an item during digitization, and the digitized item is not always a faithful representation of its image. Digitization can capture more information from a three-dimensional object, such as a bowl of fruit, than from a two-dimensional object, such as a typed recipe on paper. In the example of the bowl of fruit I photographed for digitization, several physical and sensory attributes were transferred, including the colors of the fruit, the relative sizes of the fruit pieces and the bowl, and even the texture of some fruit pieces. Compared with the digitized recipe, fewer physical and sensory attributes were transferred to this bowl of fruit. It is impossible to know how a recipe tastes, smells, weighs, or feels. The video versions of these items were more accurate in their representation, allowing viewers to see all sides of the items. Videotaping some items even provided the physical attributes of sound and scale. This process, however, was considerably time-consuming.

Working with digitized representations affects how we understand different kinds of items and our ability to use them for various purposes across several key areas. First, we need to recognize that the digitized images we access online are not always faithful representations of the original image. Second, we need to understand that the person who captured images or digitized items may or may not have been guided by standards in the process, as exemplified by the variations of the digitized image “Afro-American Army Teamsters” (Conway, 2009).

Third, we need to understand that digitized items are subjective and shaped by the digitizer. Conway (2009) argues that this is mediated by the relationship between the object’s maker and its viewer. Moreover, before digitization, the digitizer must determine whether the appearance is ideal for “certain viewing conditions”. For example, the Frey/Reilly Model shows how choices in the scanning process, based on the digitizer’s intent, can shape our understanding of a digitized image (Conway, 2009). For example, if the digitizer intends to represent an image as it appears, as in the “Ute Family” photo, the image is rendered As IS (Conway, 2009).

Fourth, we should carefully consider our interpretation of digitized items, as their meanings can be affected by the digitization process, from creating an archival master image to accessing derivatives. Image post-processing is a complex process, often with unexpressed goals (Conway, 2009).

Finally, the machinery of DH project funding constrains the range of DH projects funded, the digitized materials accessible, and the diversity of voices represented in DH collections (Terras, 2012; Terras, 2022). The DH agenda shift, increasingly driven by the economic and political landscape and by the desire for audience engagement (Terras, 2012). Consequently, funding of DH projects shapes and is shaped by visible and invisible agendas. As a result, for every digitized representation one chooses, one should consider which items are not digitized that might have provided a greater, more contextualized understanding of the image (Conway, 2009).

A Definition of Digital Humanities

Digital Humanities (DH) is an intentional interdisciplinary fusion of the humanities with advanced research and digital tools. The goal of DH is to analyze, capture, create, disseminate, enrich, interpret, store, and amplify knowledge and insights to human problems and experiences of the human condition. This field should foster collaboration among a diverse group of practitioners and scholars from various academic disciplines, who acknowledge “issues of access and intersectional power” in their work. This collaboration aims to generate, build, disseminate, evaluate, and interpret their work within the DH community and the public for engagement. DH applies current technological applications and computational methods including digital archives, digital mapping, digital archeology, data mining, timeline mapping, and network analysis, emerging technologies, such as the Shared Canvas Data model, along with traditional humanities methodologies like oral histories and the Text Encoding Initiative to enhance research, interpretation, teaching, publishing, scholarly communication, and the overall integration and exploration of technologies in the field.

My definition of DH emerged from readings by Alvarado, Booth and Posner, Drucker, Kirschenbaum, Owens, Risam, Unsworth, and others, and by examining exemplar DH projects including What America Ate, The Shelley-Goodwin Archive, Photogrammar, and Constructing the Sacred, which applied a range of DH tools in practice and reflect the spectrum of DH’s tool. My definition also acknowledges the shifting definitions of DHs over time and how these shifts highlight the dynamic and evolving nature of the field, which shifted from deciding who fits in the DH camp “or under the DH umbrella” (Bobley, 2008; Ide & Mylonam, 2004; Presner, 2010) and DH’s relationship with technology, which starts as limited (Henry, 1993; Unsworth, 2006) to more intentional and complex (Cohen et al., 2008; Schnapp & Presner, 2009). The power of DH tools is evident in the highlighted projects, whose applications cover the full extent of digital and computational technologies.

Kirschenbaum highlights the “intersecting” relationship between humanities and technology, which should be intentionally designed, as reflected in what I identify as “intentional interdisciplinary fusion.” Additionally, I emphasize the need for those in DHs to engage more directly with one another and the public, as Owens promotes. My definition also acknowledges DH’s shifting emphasis from the digital technology used to focusing on the humanities of the digital (Roth, 2019). Finally, my definition recognizes the need for DHs to concern themselves “deeply” with intersectional access and to be most valuable when they promote public engagement, humanistic knowledge, and understanding (Booth & Posner, 2020).