Wednesday, November 7, 2012

LIFE Online Photo Collection

In this post we will be examining two different implementations of LIFE Magazine’s photograph collection.  Google has implemented one version of the collection and Getty Images is hosting the second.  It is presumed that both implementations have the same photographic material, because it is possible to find certain images on both sites.  However that is difficult to verify without directly contacting each hosting organization.  Each company has their own approach to displaying the metadata associated with a photograph.  Let us jump in and examine how that information is viewed and accessed.

LIFE photo archive hosted by Google
Brook Shields showed up in the search for 1920s photographs

LIFE photo archive’s home page hosted by Google, has a very simplistic setup.  At first look it seems as though the collection utilizes two method of displaying its content.  On the left side of the home page there is a listing of decades and on the right side there is a list of categories that is subdivided into subjects.  Upon closer examination the decade listing is not always set to search the specific decade stated.  For example when selecting the 1920’s instead of getting images from the 1920’s the user is presented with images about Charles Lindbergh.  Since the links from the home page is just a search filter using Google’s technology it is possible to modify the search to only present results from the 1920’s.  This is done by entering the search term “1920 source:life”.  However when the search results displayed, we find an image of Brook Shields taken in 1994 within the first ten results.  This puts into question the metadata being used by Google.  Selecting that image to examine the record associated with it paints a not so rosy image.  For the previous example of the Brook Shields image, two elements are listed.  They are the size and the title.  The title itself combines the year and the name of the subject in the photograph.  That information should be parsed out into their respective fields.  It is unknown to the user why the photograph was brought up when searching for images from 1920. 

Another observation about the descriptive metadata is that they are being made in regards to both the digital surrogate and the content of the photograph.   There are about three basic elements of describing a photograph that a consistently listed in the sampled records.  The three elements are the identity, content and structure, and related materials.  Creator and context, location, and date are on some images while not on others.  It is uncertain whether this inconsistency between photographs is because Google does not have the appropriate metadata, or if it is because they choose not to display that information.  

Below is a list of additional problems noted about the metadata:
·         there is no statement that informs the user which controlled vocabulary, if any was used to create this online database 
·         abbreviated terms are used in the description
·         there does not appear to be any content standard or best practice guidelines being used for the metadata creation
o   an example is there is no consistency in title field, at times the title contains date information, other times it is used as the description of the photograph
 Examine the descriptions in the links below:
·         The about-ness of the photograph is skewed. The following image's description talks more about the movie being play rather than it is depicting a drive-in theater.
·        Iconic photographs are impossible to find.  An example is looking at the home page you can see the thumbnail in the 1930's section of  Dorothea Lange's Migrant mother, but clicking on the link brings up 2 images which are not of the migrant mother.  Searching Dorothea Lange or Migrant mother also does not return the desired photograph.

The online collection setup by Google is still very text oriented and does not work well with photographic material especially when the metadata it uses to perform the searches are not well formed. 

LIFE online collection at GettyImages
GettyImages implantation overall is very useful and provides enough detail to make a record easier to find and access.  The main page of the collection displays thumbnails representing categories that the photographs have been divided into.  GettyImages has not created a category level description for each section that they created.  However after selecting a category to browse, the photographs can immediately be refined by specific event, people, keyword, or photographer.  This gives the user a quick idea of the what, who, where and when of the section.

When examining the details of the records we can see that five basic elements for a photographic record are consistently listed in a record.  They are the creator and context, identity, content and structure, acquisition and appraisal, and general notes.  When looking for related materials the GettyImages implementation is lacking.  The only instance when related materials can be found is when browsing.  If a user enters a search and selects a photograph there is nothing that states the category of the found record.  There are keywords and subject access terms listed in the record, but they are not linked to other images.  While on the top of subject access, there is no authority control mentioned.  One can only presume that the system is using AAT or TGN since this is hosted by GettyImages, but that information is not explicitly noted.

On the left the keywords are delimited.  On the right the keywords are not.
Very similar to the use of an authority control, there seems to be a content standard or best practice guideline being followed.  However that standard is not noted and it does not seem to be followed strictly.  An example is that some records the keywords are delimited with a comma, while other are listed one after another with only a space separating words, so it would be difficult to pick out subject terms.

The record describes a mix of the digital surrogate and the content of the photograph.  This can be seen in the records when you notice that size and resolution information is given about the photograph, but then the caption describes the about-ness of the image.  This is most likely because both the Google and GettyImages implementations are setup to sell the photograph and having size information is important to a potential buyer. 

Example of a Subjective description
The caption is used as a descriptive field.  This is problematic on many levels with this implementation.  The first issue is that the title field uses a truncated version of the caption and therefore not very useful as a title.  The other issue is that the caption is subjective.  In the following image the caption states that it is depicting a man modeling liquid crystal thermography.  It then goes into an interpretation of the photograph.  This is not always the case, but in several of the photographs examined, a subjective view was noticed.  A hypothesis would be that the captions were created by the photographer and transcribed into the record.  An indication that points to the validity of this hypothesis is that the title field is just a truncated version of the caption, which one can conclude that either this was prescribed by the content standard, or a means of saving time when entering the data for the record.  Another evidential point is that for images that can be found in both Google and GettyImages, the caption is exactly the same.  This shows a shared metadata source between the two.  It is also important to note that Google has not harvested GettyImages’ metadata, because the image dimensions are different.  This shows that those captions are somehow associated with the physical photograph.

When examining the two implementations of the LIFE online photo collection, we can see that both have their advantages and disadvantages.  If you are looking for information on the photograph then you should go to GettyImages.  If you are looking for lots of similar images, you should use Google.  Overall a recommendation of user friendly access would go to GettyImages.  This is because of their richer metadata and organizational structure that allows users to better search their collection, and therefore having better access to the collection.

No comments:

Post a Comment