-
Notifications
You must be signed in to change notification settings - Fork 3
The Purpose of DM
From the Introduction of Tim Andres's undergraduate honors thesis on DM
With the explosive popularity of cloud based document storage, even novice computer users have become familiar with the enormous benefits of storing their work online. Services like Google Docs allow users not only to store their work online, but also to easily collaborate with others in real time. These sorts of services usually also allow users to make their work available to others by sharing it with specific individuals directly, or making it publicly accessible and searchable.
Services such as Google Docs, Microsoft’s Office 365, and Apple’s iWork for iCloud already do an admirable job of making these sorts of documents available for collaboration and sharing, but they were clearly designed with the modern office in mind. This means that they are inherently limited to word processing documents, presentations, spreadsheets and the like. While this may work nicely for creating new content, and perhaps importing a decade or two of digital productivity documents, it leaves out an unfathomable mass of documents which were created in the days before modern computing using pen and paper, quill and canvas, and even chisel and stone. These documents in their physical forms hold a wealth of information beyond merely the words to be transcribed into a text file. While many scholars may study the transcribable content of these documents, others may be more interested in examining the handwriting of the scribes to learn about the people who wrote the documents. Others still may be interested in using visible and hidden signs of editing, or even features as simple as notes scrawled in a margin, to infer the thought process of the original writer. In order to make these documents available in a digital format without losing this embedded information, clearly something very different from a traditional word processor is needed.
Thanks to the power of digital photography and ever decreasing cost of digital storage, many traditional repositories of these physical documents have been digitizing these resources as archival quality image files and making them available online. They are sometimes aggregated with specialized images of the documents, such as photographs taken under specific wavelengths of light to reveal hidden writing within the original document, and then assembled into a sequence of images which represent the original document. In the past, the data which connected these images in their proper order with associated metadata was arbitrarily structured by the organization responsible for archiving the documents. The images would then be displayed in some sort of bare bones image viewer, usually just a webpage with forward and backward buttons and perhaps a table of contents. The lack of a standardized data model has acted as a stumbling block to sharing the documents with software used by other organizations to store and view those documents, which in turn made it more difficult to develop a single software tool to manipulate documents from different sources.
Perhaps even more critically, in the past, there has been no standardized model for representing the data generated by scholars studying these documents, hindering the free flow of academic knowledge. The reported common practice among medievalists, for example, has been to use basic image manipulation software to crop sections of interest out of larger images, perhaps draw markings to indicate features of note, and copy those images into a word processing document with an arbitrary notes structure. Others have even come up with methods involving the use of Excel spreadsheets to track their work. Not only does this sort of system fail to maintain easily followable bi-directional connections between regions of interest and scholarly annotations; it fails to allow other scholars to connect and share their own work on the same documents. Multiple groups could easily spend months working on the same documents independently, only to perhaps eventually stumble upon each other’s work after it has been published. Clearly, this is all far from ideal. One would want to see the majority of archived documents represented in a standardized format so that any software which supported the standard could utilize archived resources from a variety of sources. One would also like to see a standardized model for representing the work of scholars on these resources – flexible enough to be used by a variety of disciplines, yet well defined enough that the data is still processable for purposes like searching.
While data standards are crucial to the creation and sharing of this information, certainly no one expects the majority of these scholars to write out XML files by hand in order to share their work (although such attempts are not unheard of). An intuitive user interface is essential to facilitating the creation, sharing, and publishing of this data. The DM Project was created with this purpose in mind in 2008, building a web based user interface for scholars to view and annotate manuscripts, canvases, and transcribed text documents, along with a supporting back end server to centrally store and serve the generated data. Originally called the Digital Mappaemundi project from the latin for “maps of the world,” the project gradually expanded to include more generalized resources like manuscripts, and its name was shortened to “DM”. At this time, the project currently has over sixty humanities scholars as beta users, all generating data which can be exported in standardized formats to be shared with the world. These scholars have already used DM to annotate maps, scrolls, manuscripts, and even images of archaeological digs, and they are currently sharing their work through DM with their colleagues, students, and readers.
While the DM Project has been designed with the use case of medieval scholarship in mind, there is nothing in the interface of DM nor the data models it relies upon which restricts its use to those disciplines. In fact, the annotation tools which DM provides could be of use to practically any discipline whose practice involves the study of some sort of text-based or image-based data. Just as easily as DM can be used as a tool for medieval scholars to annotate ancient manuscripts, it could be used by a biologist to annotate microscopic images of cell cultures, a geologist to mark up images of strata, or even a medical doctor to make a few notes on an X-ray image. Each of these disciplines would likely benefit from some specialization of the interface and data to handle their specific use cases, but both DM and the data models it uses have been built with extensibility in mind. Furthermore, DM is being released as open source software, meaning that any developer can build upon DM’s extensive codebase to build software which fits a user group’s needs, and any institution can set up its own private instance of the tool on its own servers.