Chris and I have begun auto cropping large batches of images. Photoshop CS3 does have an intriguing auto crop command, but it does work effectively enough for our purposes, typically only resizing part of the image. Chris has experimented with auto cropping images divided by even and odd (left and right bound pages of the newspaper demarcated by year and month), and found it effective when compared to the time taken to crop them by hand and convert them from tiff to jpeg one by one. Most of the time is now spent fixing images where text is cut off by the auto cropping action. The result isn't quite as precise, but for this specific project, they don't have to be. We've found that by auto cropping images we can sometimes finish in about 2 hours one year's worth of newspapers as opposed to 4 hours when manually cropping (although we sacrifice some quality). So what we have is a workable solution for the rest of the semester (why didn't we think of this earlier?!), but a long term investigation of more sophisticated software or more precise means of cropping is still worth looking into. Last week, we also corrected a number of images that did not have banding for the Daily Iowan. Content DM makes correcting such errors pretty simple. We simply recalled the images after changing the options, reuploaded and approved the images.
This past week, I also attended two of the David Eads presentation on Drupal, a content management software. Most of the attendees seemed pretty firmly entrenched in IT, tossing around sophisticated technical questions with aplomb, while I often had no clue what they were talking about. But while I may not have garnered a lot from the technical explanations of the software, I thought his explanation of the implications of open source software was intriguing (I love discussions of upcoming trends and their subsequent effects on society). I find it fascinating that the internet, originally intended as a tool for collaboration, has in a sense, vacillated back and forth on this idea of community and what defines a community. The internet seems destined to forever remain in large part proprietary in nature, and yet, wikis and enormous communities of users, like the ones that use Drupal, continue to bring disparate groups of people together. The larger social implications of people (are we naturally inclined to do so?) working together are (although not unusual) quite profound. Now we have the technology to make communication and publishing easier than ever.
Some of the advice doled out by Eads included:
-Pick a vendor you like rather than a technology (emphasizing importance on working with the right people)
-Folksonomies are a great example of leveraging technology in a creative way. Another example might be the naming game Google uses to disambiguate images
-Change the mindset of I'm going to be everything to everybody. It's better to be something to somebody.
-In the future of the web may be in the Semantic Web or getting better metadata about various pieces on your page.
-The web will become a more seamless media experience (on this I wasn't quite clear what he meant. In the days of the nascent internet people seemed to believe the next thing would be something better than TV. Well we now know the internet is nothing like TV. So will the next big thing be something like the internet, but better or more seamless? Professor Hsieh discussed the implications of intuitive touch screen technology in computing. Will the next big thing be something like the iphone? Internet that is portable?).
We seem to be nearing the end now with little over a month left to go on the DI and Glenister project. Chris and I met with Nicki Saylor and Mark Anderson this week to go over some specifics about how best to begin wrapping up the projects with the long-term future of the projects in mind. Nicki proposed focusing on the following goals:
1)Create an interface for both collections--This should be fairly straightforward as Nicki wants a largely uniform layout for each of the digital library collections. So it's mostly a matter of selecting images and subjects for canned searching (although there's not a lot of variety of searches that could be constructed for the Daily Iowan).
2)Ramp up the workflow--Chris and I are to look at ways to possibly expedite the process of batch cropping and converting the images (from Tiff to jpegs). A large part of our time this semester has been devoted to (sometimes mindless) labor as opposed to intellectual work. Part of what we discussed in seminar was trying to find a balance between sticking to a familiar process of work or risking trying something new that may in end save you time in the long run, or perhaps not. The nice part about being a digital fellow, is that our schedules are flexible enough that we don't have to become overly preoccupied with deadlines and even when mistakes are made, they become valuable learning experiences. This I feel must be part of the "real world" work experience because while achieving short term immediate goals are important (justifying budgets commensurate with productivity) experimentation are the keys to long term progress (in other words allowing time to play!). Chris and I will look into possibly finding software or some kind of batch command for photoshop. Nicki stressed finding a workable solution, one that may be imperfect but still help ameliorate some of the labor intensive work that has been so far involved.
3)Nicki is going to draft a memorandum of understanding for the Glenister project. Chris and I will edit it to the best of our knowledge and then it will be passed on to Tiffany Adrain.
On a non-digital project related note, I attended the Competitive Intelligence conference held on Friday. While I'm not sure I would like to go into the business information field, the conference did give me a sense of how important research skills are to competitive companies. One of the most interesting pieces of advice that was given at the conference was for librarians or library students to join professional organizations that don't have "our" skills. My previous belief was that the general public, large companies included, lack an understanding of the importance of the work of librarians, but perhaps this is not as pervasive as I once thought.
I had a conversation with fellow fellow Joe the other day about the progress of our projects and I think we're all starting to realize just how little we're going to finish this term. Relatively speaking. This of course was as I was coming into the semester with high expectations and almost zero understanding of how to begin a new project. 6,000 slides did seem like a large undertaking, but I figured maybe, if I worked assiduously enough, I could finish like half of the slides? As it stands, we have some 120 slides uploaded on the digital library site with another 100 ready to go. And we're more than half way through the semester. Unfortunately, it's not a matter of, well maybe we can just pick up the pace. It's been about two weeks since the Geoscience department last emailed me, and I felt like Chris and I had to do a bit of haggling to get our last batch. I don't mean to sound overly negative about this aspect of the project. I understand the Geoscience Repository employees are busy with a multitude of other things of more pressing concern. Moreover, Tiffany Adrain is dependent on Dr. Glenister freely giving of his own time to, somewhat, laboriously, go through and label the collection, slide by slide.
Although, we've reached the mid-1880s in the Daily Iowan collection, we haven't been regularly uploading the images because of a problem with Content DM. Mark promptly emailed OCLC, and we found out, today, that the problem had to do with spacing the cursor after the final line in our tab delimited text files. To which I say, why is Content DM so finicky? and I never ever would have thought that was the problem. Well, at least, a simple problem only demands a simple solution. Hopefully, we'll get at least another decade posted to the collection site in the next week or so.
I feel lucky that I've had so many helping hands in my own projects in the DLS department. Chris and I have encountered our share snafus, and Mark has repeatedly stressed that it is all a good learning experience. And that, I feel, is what I will probably take away from this semester. Yes, it would be nice to have some evidence for all the work we've put into metadata and such, but the more salient, albeit less explicit, goal is to establish a workflow and create documentation for our standards to keep our project running so that maybe a year or two down the line, someone else can finish the work we started.
Finally, in a non-digital, but library related note, I went to the ILA conference on Friday and found it, a bit to my surprise, to be pretty worthwhile. While, many of my fellow library students seem much more self-assured about the direction in which they're going (i.e. they want to work in academic libraries or with old manuscripts in special collections or in school libraries), I still feel like I've been mostly direction-less since beginning school. Part of me wonders if I entered library school as an alternative to doing something more challenging or perhaps less job-friendly. Did I want to be libarian because it was something that is comfortably familiar? While much of what was discussed at the conference, wasn't earth-shatteringly new, perhaps it was not precisely what was said, but how it was said (Personally, I felt quite inspired). Speaking in broad terms, Joseph Janes, a speaker from the University of Washington, was able to precisely articulate why the library profession is so worthy an endeavor. The idea that libraries extend beyond a building isn't new, but as Janes put it, libraries have always tried to get out of the building, and it is the digital world (as David Weinberger puts it, the conversion of atoms to bits) that allows the library to be somewhere and everywhere. But in doing so libraries shouldn't strive to be Google. Libraries can beat (maybe an imprecise term?) Google by not trying to be Google, but focusing on the niche areas where our quality will always trump the expedient, yet shallow search results of Google. The digitization of documents by the staff at DLS, and other similar departments at Iowa, are working towards ensuring libraries stay relevant in an ever more digital world.
At Professor Srinivasan's behest, Chris and I looked at a number of other digital newspaper collections this week, for comparison's sake and hoping to glean some useful ideas for our own. One particularly impressive example is the Yale Daily News Digital Archive. One of the cool and useful features in this particular collection is the ability to highlight and select segmented text for close viewing. According to Mark Anderson, this feature is only available through OCLC (for a considerable fee) employing a mapping standard called METS/ALTO, which delineates each articles position on a specific page. More can be read about OCLC's proprietary control over this feature here.
I'm feeling much more confident about where I stand with my DLS work, or at least better than I was feeling about a week ago. On Monday, Chris and I met with Tiffany Adrain again to discuss our metadata schema. Unfortunately, we weren't able to meet up with Dr. Glenister, who apparently only drops in sporadically to meet with Tiffany. We also finished uploading the first batch of Glenister slides to the Digital Library website. It was really great seeing what the images look like on the site, but perhaps more importantly, it helped me wrap my mind around the sheer enormity of our task. Only 6,900 more slides to go.
Chris and I also talked to Mark about the possibility of starting another project in our downtime, while we wait for Tiffany to send us more metadata. Tiffany had explained she was still in the process of hiring another graduate student to work exclusively on the Glenister slides. Moreover, many of the slides require the tacit expert knowledge that only someone like Tiffany can provide. Thus, what will probably be a regular delay in receiving metadata information. Mark suggested working on the Daily Iowan project, which Chris had worked on over the summer and which comprises over 3 Terra Bytes of digitally scanned newspapers (in Tiff format) dating from the mid-nineteenth century. I'm pretty excited about getting started on this as it is a bit more contiguous to my own interests (I wrote for the student newspaper for some 8 years in high school and college). To begin, we've focused most of our time on cropping the images and changing the format to accommodate OCR (optical character recognition). Chris has been really patient in teaching me the applications for the Daily Iowan project and familiarizing me with the work flow. I'm looking forward to a productive week!
Chris and I were scheduled to met with Tiffany Adrain again this morning to touch base about where we were in terms of metadata. As Mark explained it last week, our job is to look at the metadata in more broad terms, given our lack of geoscience knowledge. Over the course of last week we filled in the metadata categories the best we could, given we were working with a limited number of defined categories like the original labels written by Glenister and some additional classification made explicit by Ms. Adrain. We also made a trip to archives last week to see if we could dig up some background information on Glenister's research, perhaps in the form of articles or training notes. According to David McCartney, university archivist, special collections has some 27 square feet of boxes filled with Glenister material that has yet to be cataloged. Going through the first box of material gave me a better appreciation for the detph and considerable impact Dr. Glenister has had in his respective field. Additionally, the man is seriously well traveled, as indicated by the correspondence that is documented in some of his archive material. I do hope I have a chance to meet with him at some point, if only to get a better sense of who he is. This week Chris and I planning to finish uploading the first batch of slides. Depending on the lag time between finishing that job and when Tiffany sends us more metadata, we may talk to Mark about starting or helping out on another job.
Chris and I made progress on setting up the Geosciences collection. On Tuesday, we met with Mark to go over, to the best of his knowledge, what the Geoscience project would entail. Chris introduced me to some of the basics of migrating images into ContentDM. On Wednesday, we met with Tiffany Adrain who filled us in on some of the history of the slides and Brian Glenister, who has, by all accounts, led a colorful and illustrious life entrenched in the field of geosciences. Tiffany explained that many of the slides were labeled in a fairly non-descript, sometimes arbitrary, fashion. The labels written on the slides by Glenister, such as Category A or B or sometimes a long list of numbers, may or may not have meaning to Glenister, but appear to have been written for purely organizational purposes. Or more exactly put, divided into categories that perhaps only Glenister or a seasoned geologist might understand. Most of the 7,000 have been scanned (a small proportion are attribued to another photographer, but have been approved to be included in the collection) and tagged by the labels Glenister provided. Mark and Tiffany discussed which categories of metadata might be most appropriate, although considering the size of the collection, expediency might take some precedence over quality. Tiffany said she was in the process of hiring a student to assist with the development of more detailed (and hopefully more meaningful) metadata and was planning to set up regular (possibly weekly) meetings to begin going over the slides with Glenister. In the meantime, Chris and I began the process of automating the resizing of images for the collection. On Friday, we met with Ellen, who in lieu of metadata librarian Jen Wolfe, gave us some standardized metadata categories to work with. Tiffany emailed us back on Friday, asking us if we were free to meet with Dr. Glenister on Wednesday. Despite having only scratched the surfaced of this no doubt considerable collection, I'm really looking foward to meeting the man behind the photos. I imagine he has some interesting stories to tell.
