Article discussed: Alexander, Arden and Tracy Meehleib (2001). "The Thesaurus for Graphic Materials: Its History, Use, and Future." Cataloging & Classification Quarterly 31(3/4): 189-212.
February's Metadata Discussion Group session was a lively one. The topic of subject vocabularies beyond LCSH sparked a great deal of interest. The session began with discussion of why a separate subject vocabulary for graphic materials was needed, especially in the Library of Congress. Some participants had even cataloged pictures or posters with LCSH, not knowing that other options existed. Participants realized the need for subject terms that were not in LCSH for describing photographic materials, but recognized the potential to add these terms to LCSH rather than starting a new subject vocabulary. The primary reason for needing a separation of subject vocabularies identified during this discussion was a difference in the level of specificity needed for cataloging visual material as opposed to textual material.
Participants then noted that LCSH and TGM I are structured differently; LCSH is a subject heading list while TGM I is a true thesaurus. While this is an important distinction to understand, the group was uncertain as to the specific implications for practice. Both are standardized vocabularies and are applied in a similar fashion. In the last 15 years LCSH has become more thesaurus-like in standardizing cross-reference structure and describing narrower, broader, and related terms instead of see and see also references.
Overall the discussion group thought that the existence of TGM has struck a reasonable balance between one big general vocabulary and lots of little specific ones. While TGM is specifically focused on graphic images, that is a big space and TGM can be applied in many ways. For a big image collection, a graphic materials-specific vocabulary is a great deal more useful than LCSH would be. The group expected image cataloging (and TGM use) to continue to grow as libraries focus more and more on special collections.
From here, the discussion moved on briefly to comparing the top-down design approach of TGM II with the bottom-up (literary warrant) approach of TGM I. A significant issue with the bottom-up approach was identified - that it is difficult and time-consuming to maintain a robust reference structure for a vocabulary that is constantly growing.
The topic of whether or not to subdivide TGM was a main focus of this month's discussion. A participant noted that the precoordinated approach has its origins in the printed card catalog, where it was necessary. Now that we are in online systems, this approach can be rethought.The subdivided approach takes more time to apply (this was consensus but nobody knew of data to cite) and it's not possible to be as specific with geographic locations in subdivisions than it is with postcoordinated geographic headings. Postcoordinated approaches allow the user to decide which feature is of primary interest, rather than having one selected ahead of time. Subdivisions also introduce redundancy as the same subdivisions are often applied to many main headings. But are there cases when they would be different? Would a TGM I heading ever have a different time period subdivision than a TGM II heading on the same record? Perhaps in the case of a contemporary poster of a historic event? It would be more difficult to make this distinction in a postcoordinated approach. A potential benefit of a precoordinated (subdivided) approach is the creation of a browse index. This is achieved in a different way via faceted browsing with the postcoordinated approach. The group felt strongly that the most important goal was to produce a product that is easy and understandable for our usrs. More user studies are needed to learn more about this issue.
The group then wondered what the literature on precoordinated vs. postcoordinated vocabularies looks like. Is there anything recent? Thomas Mann wrote recently on this topic, but no one was aware offhand of other recent work other than Lois Chan describing FAST.
At this point, the discussion turned to the type of training that would be necessary for someone to effectively apply the TGM. For TGM II (genre), the individual would need some level of background with formats of graphic materials. But for topic, participants thought that the same training to perform subject analysis on textual works would apply to graphical works. For some image materials, it is necessary to become familiar with important buildings and people likely to be in the collection, for example buildings on the IU campus and IU presidents for photographs in the University Archives.
The discussion wrapped up with thoughts on the lack of information inherent in the resource that helps with the cataloging process for graphic materials as opposed to textual materials. Generally images come with something that helps identify the content and its origin. Given at least a small amount of information, a cataloger would apply the same type of research techniques, including those applied for authority work, that are already in place in many cataloging units. Image description could be portrayed as an extension of existing work rather than a departure.
Saturday, March 14, 2009
Monday, February 2, 2009
Summary of MDG Session, 1-29-09
Article read: Chris Freeland, Martin Kalfatovic, Jay Paige, and Marc Crozier. (December 2008). "Geocoding LCSH in the Biodiversity Heritage Library." The Code4Lib Journal 5. http://journal.code4lib.org/articles/52
As with many MDG sessions, this one began with a discussion of unfamiliar terminology in the article read. This article contained a few technical terms that members were interested in hearing more about, likely due to the fact that the audience for the Code4Lib Journal is primarily programmers rather than catalogers or metadata specialists. One participant wondered where the term "folksonomy" came from. Nobody had an answer, although some thought it had been around a decade or so, and Clay Shirky's "Ontology is Overrated" was mentioned. (The Wikipedia article on folksonomy credits the term to Thomas Vander Wal in the early 2000's.)
The substance of this month's discussion began by addressing the question: Would you catalog differently if you knew the data were to be used in this way? Participants noted that the burden is on the cataloger to verify and provide information that isn't immediately obvious from the resource itself. The limits of MARC/AACR2 practice (missing geographic headings in some cases) described in the article are very real – if the terms aren't there you can’t build this type of service. If you know the data is going to be used in this way then you make more of an effort to provide it. Participants repeated an often-heard comment about MARC cataloging - that populating the fixed fields takes a great deal of effort, but few systems use them. This discourages catalogers from populating them, which discourages system designers from using them... The current environment doesn't make it easy to justify doing the work to create the structured data that's really needed to provide advanced servcies.
The conversation then turned to where geographic data to support a service like the one described in this article would be in a MARC record, if those records were created with this use in mind. One important point to note is that the level of specificity is different between the coded geographic values (043, country of publication in fixed fields) and what is present in LCSH subdivisions. The former are generally continent/country/state level, while the latter can be much more specific. Discussion of these fields led participants to note that these fields represent different things - the place something is published is of interest in different circumstances than the place something is about. This represents one area (of many) where system designers need to have an in depth understanding of the data. Building a resource with more consistent geographic data (say, always at the state level) would alleviate some of the challenges described in this article, but leave out users who are interested in more granular information than an implementation like this could provide.
Some participants advocated that to promote services like the one described in this article, one should use a vocabulary that's designed specifically for geographic access only for this purpose, such as the Getty Thesaurus of Geographic Names or GeoNet. One advantage of these types of vocabularies is that they are based on "official" data of some sort (for example, the US government, the UN), whereas LCSH is based on literary warrant. LCSH therefore might not match up well with current and detailed places such as those one would ideally want for a resource map interface. Similarly, AACR2 treats some objects with geographic features (for example, buildings) as corporate bodies, which are subject to different rules for cross-references and the like.
Participants noted that there have been significat successes in geographic and user-friendly access in the MARC/AACR2/LCSH stack, however. MARC records for newspapers have a 752 field with semi-structured data listing country-state-county-city. The terms used in this field come from the authority file. The 752 field represents an early example of a field existing in response to user discovery needs. Would it be possible for us to generate this data automatically for other types of resources?
The conversation at this point moved to user behavior in general. A participant noted that at the
recent PCC meeting at ALA Midwinter, Karen Calhoun gave a presenation describing OCLC's recent research on user behavior. Their conclusion is that delivery is becoming more important than discovery. Does this mean libraries should start changing our priorities?
Different types of discovery were then briefly discussed, noting that one wants different things at different times. The subject search serves a different purpose than the keyword search. Especially for scholars, the former is useful for introductory and overview work. When delving deeper, looking for the obscure reference that will serve as a key piece of original research, the latter will be more useful. The "20% rule" for subject cataloging is one reason for this. Are tag clouds of subject headings therefore useful? Participants thought they were for some types of discovery. Other possibilities would be clouds of Table of Contents data and full text. All would have different uses, and for some the cloud presentation might be more effective than others.
A significant proportion of the discussion in the second half of the session revolved around ways to integrate together different types of geographic access. The first topic on this theme was one of granularity - how specific should the geographic heading be? Why shouldn't we provide acces to a famous neighborhood in a big city? Using the structure of a robust geographic vocabulary as part of a discovery system would help with this issue.
The changing of place names and agreed-upon boundaries over time was raised next. A place with a single name might have different bondaries over time. Political change is ongoing, and one place does not simply turn into another; maps are constantly reorganizing. Curated vocabularies such as LCSH and TGN take time to respond to these changes. Is it necessary to update older records when place names change? Participants settled on the standard answer: it depends. For resources such as biological specimens, current place names are likely to be more useful, to assist the researcher with understanding the relationships between them over time. For works about specific places, the place as it existed during the time described is more important.
The next issue raised in the session was that geographic places don't exist in a strict hierarchy. National parks, rivers, and lakes, for example, aren’t within single states. LCSH headings exist for these, and for rivers can be separate for each state the river crosses. Participants were not certain if cross-references existed between the river name and the headings for all states it crossed, which would create a machine-readable link between the two.
It was at this point that GIS technology as a solution came up. By defining everything as a polygon rather than a label with some classification of type ("state," "river," park"), geometry can be used to retrieve places relevant to a specific point. Effectively connecting all of these overlapping but not exclusive things in traditional library authority files would be a challenge. Many other geographic-type units could be used for retrieval, including zip codes, area codes, and congressional districts. These change over time as well, further coplicating the situation.
The final issue raised in connection with geographic access was the notion of places being referred to with different names in different languages. Libraries are increasingly adding cross-references from multiple scripts and languages into authority files. This is a good thing, certainly. The lack of a 1:1 mapping from historic places makes this difficult. Even for the residents of a place, the dominant language changes over time and therefore the "official" name.
The Virtual International Authority File is attempting to address this issue by linking together names for the same places from multiple national authority files. It's a bit unclear what the status of this project is, though. LC and OCLC consistently report progress but no clear indication of when it's going to become a production system.
As with many MDG sessions, this one began with a discussion of unfamiliar terminology in the article read. This article contained a few technical terms that members were interested in hearing more about, likely due to the fact that the audience for the Code4Lib Journal is primarily programmers rather than catalogers or metadata specialists. One participant wondered where the term "folksonomy" came from. Nobody had an answer, although some thought it had been around a decade or so, and Clay Shirky's "Ontology is Overrated" was mentioned. (The Wikipedia article on folksonomy credits the term to Thomas Vander Wal in the early 2000's.)
The substance of this month's discussion began by addressing the question: Would you catalog differently if you knew the data were to be used in this way? Participants noted that the burden is on the cataloger to verify and provide information that isn't immediately obvious from the resource itself. The limits of MARC/AACR2 practice (missing geographic headings in some cases) described in the article are very real – if the terms aren't there you can’t build this type of service. If you know the data is going to be used in this way then you make more of an effort to provide it. Participants repeated an often-heard comment about MARC cataloging - that populating the fixed fields takes a great deal of effort, but few systems use them. This discourages catalogers from populating them, which discourages system designers from using them... The current environment doesn't make it easy to justify doing the work to create the structured data that's really needed to provide advanced servcies.
The conversation then turned to where geographic data to support a service like the one described in this article would be in a MARC record, if those records were created with this use in mind. One important point to note is that the level of specificity is different between the coded geographic values (043, country of publication in fixed fields) and what is present in LCSH subdivisions. The former are generally continent/country/state level, while the latter can be much more specific. Discussion of these fields led participants to note that these fields represent different things - the place something is published is of interest in different circumstances than the place something is about. This represents one area (of many) where system designers need to have an in depth understanding of the data. Building a resource with more consistent geographic data (say, always at the state level) would alleviate some of the challenges described in this article, but leave out users who are interested in more granular information than an implementation like this could provide.
Some participants advocated that to promote services like the one described in this article, one should use a vocabulary that's designed specifically for geographic access only for this purpose, such as the Getty Thesaurus of Geographic Names or GeoNet. One advantage of these types of vocabularies is that they are based on "official" data of some sort (for example, the US government, the UN), whereas LCSH is based on literary warrant. LCSH therefore might not match up well with current and detailed places such as those one would ideally want for a resource map interface. Similarly, AACR2 treats some objects with geographic features (for example, buildings) as corporate bodies, which are subject to different rules for cross-references and the like.
Participants noted that there have been significat successes in geographic and user-friendly access in the MARC/AACR2/LCSH stack, however. MARC records for newspapers have a 752 field with semi-structured data listing country-state-county-city. The terms used in this field come from the authority file. The 752 field represents an early example of a field existing in response to user discovery needs. Would it be possible for us to generate this data automatically for other types of resources?
The conversation at this point moved to user behavior in general. A participant noted that at the
recent PCC meeting at ALA Midwinter, Karen Calhoun gave a presenation describing OCLC's recent research on user behavior. Their conclusion is that delivery is becoming more important than discovery. Does this mean libraries should start changing our priorities?
Different types of discovery were then briefly discussed, noting that one wants different things at different times. The subject search serves a different purpose than the keyword search. Especially for scholars, the former is useful for introductory and overview work. When delving deeper, looking for the obscure reference that will serve as a key piece of original research, the latter will be more useful. The "20% rule" for subject cataloging is one reason for this. Are tag clouds of subject headings therefore useful? Participants thought they were for some types of discovery. Other possibilities would be clouds of Table of Contents data and full text. All would have different uses, and for some the cloud presentation might be more effective than others.
A significant proportion of the discussion in the second half of the session revolved around ways to integrate together different types of geographic access. The first topic on this theme was one of granularity - how specific should the geographic heading be? Why shouldn't we provide acces to a famous neighborhood in a big city? Using the structure of a robust geographic vocabulary as part of a discovery system would help with this issue.
The changing of place names and agreed-upon boundaries over time was raised next. A place with a single name might have different bondaries over time. Political change is ongoing, and one place does not simply turn into another; maps are constantly reorganizing. Curated vocabularies such as LCSH and TGN take time to respond to these changes. Is it necessary to update older records when place names change? Participants settled on the standard answer: it depends. For resources such as biological specimens, current place names are likely to be more useful, to assist the researcher with understanding the relationships between them over time. For works about specific places, the place as it existed during the time described is more important.
The next issue raised in the session was that geographic places don't exist in a strict hierarchy. National parks, rivers, and lakes, for example, aren’t within single states. LCSH headings exist for these, and for rivers can be separate for each state the river crosses. Participants were not certain if cross-references existed between the river name and the headings for all states it crossed, which would create a machine-readable link between the two.
It was at this point that GIS technology as a solution came up. By defining everything as a polygon rather than a label with some classification of type ("state," "river," park"), geometry can be used to retrieve places relevant to a specific point. Effectively connecting all of these overlapping but not exclusive things in traditional library authority files would be a challenge. Many other geographic-type units could be used for retrieval, including zip codes, area codes, and congressional districts. These change over time as well, further coplicating the situation.
The final issue raised in connection with geographic access was the notion of places being referred to with different names in different languages. Libraries are increasingly adding cross-references from multiple scripts and languages into authority files. This is a good thing, certainly. The lack of a 1:1 mapping from historic places makes this difficult. Even for the residents of a place, the dominant language changes over time and therefore the "official" name.
The Virtual International Authority File is attempting to address this issue by linking together names for the same places from multiple national authority files. It's a bit unclear what the status of this project is, though. LC and OCLC consistently report progress but no clear indication of when it's going to become a production system.
Wednesday, January 14, 2009
Summary of MDG Session, 12-18-08
Article discussed: Kurth, Marty, David Ruddy, and Nathan Rupp. (2004). "Repurposing MARC metadata: using digital project experience to develop a metadata management design." Library Hi Tech 22(2): 153-165.
The discussion group felt that while it was desirable that the work described in this article was based on theoretical work on metadata management, the explanation of the metadata management theory, including the concept of enterprise, was not extensive enough to fully understand the connection. It was clear, however, that to do management, you have to do mapping and transformation. Management allows you to rethink and retool. Our group was interested to know what has happened since this article was written. Have they put this into production? What has changed? It appears there is a follow-up article to this one that would be interesting to read.
The article claims that MARC mapping work is representative of the metadata management task as a whole. Choosing metadata standards based on specific project needs is good, and the projects described here demonstrate how to do that. It's easy to imagine a project where you can start with MARC. But what do you do when no MARC already exists? At IU we have experience in many library departments wth projects that re-use existing MARC metadata.
The group identified three possible cases for metadata management for a digital project: have existing MARC, have existing non-MARC, have no existing structured metadata. Are the strategies outlined in this article useful in all three of these cases? We didn't come to a strong conclusion on this issue.
An interesting discussion grew up around the topic of how to deal with legacy (pre-AACR2) MARC records? Institutional memory is likely the best bet, as documentation comparing older practices and current ones is sparse. Politcal boundaries change and places of publication become no longer correct. Some legacy data is easier to deal with, however. An institution could use an authority vendor to update name headings with death dates. Yet certain data elements should be updated over time, but others shouldn’t. The group noted that most metadata work is bibliographic record based and doesn’t do enough with authority records. Making the full authority structure available to the metadata creation staff is sorely needed.
A substantial amount of discussion time was spent on the topic of collection-specific mappings. The benefits of corse are that these get it done, the way you want it. The drawbacks are potentially reduced shareability and interoperability. One has to take the whole scope of the project in mind to make good decisions and worry about what’s really important. Have to keep USER in mind. This is difficult to do, though. We think “the user needs this information” but we should think “how can the user use this system?” One participant noted that we worry too much about the specialized discovery case to the detriment of the generalized one. How much tweaking of metadata mapping is of use? The community seems to swing back and forth over time between the generalized and specialized approaches.
The discussion then turned more theoretical, with thoughts on the changing roles of libraries – specifically, to what degree should we be the intermediary? If the user is on his or her own, should this change the way we provide access to information? We do see a great deal of evidence that libraries have moved to a model where users interact directly with information with no active intermediation from us. The system provides the intermediation that staff once did. We expect better technologies to automatically enrich our records in the future to help with this. For us, participants felt it was more important to get something out than to get it perfect. We need to make a better effort to integrate authority control into non-MARC environments. Automated methods will rely on the authority records a great deal. It therefore follows that we should send less time on bibliographic records and more on authority work. The MARC world is certainly moving in this direction, with professional catalogers doing more high-value activity, leaving the lower-value tasks to machines or lower-level staff. Mapping activities are an example of the higher-value activity, as seen in this article.
This article describes the most common transformation as MARC to simple DC. To make sure information gets into the right DC fields, one need to understand DC. Those doing the mapping must ask - what is the essential information to go in DC? What really identifies rather than just describes? The role of the cataloger would be to oversee the transformation process, to make sure it works correctly. This would need to happen both on the content end and the technical end.
What should relationship of metadata staff to technical staff be? Metadata staff understand both the source and the target data. They would still have to correct things in the output in the end. It certainly helps if the technical staff understand the data as well. Similarly, metadata staff need to have technical skills. For metadata staff, understanding non-standard source data can be a big challenge. The Bradley films are an example of these challenges here at IU. Each set of materials will have different balance of effort spent on it, based on perceived importance and use. Mapping often unearths mistakes in the original metadata. We must get the best bang for our buck by spending more time on the information that’s really important for the users, and leave the rest alone. Effective projects will also need the involvement of collection development staff.
The discussion group felt that while it was desirable that the work described in this article was based on theoretical work on metadata management, the explanation of the metadata management theory, including the concept of enterprise, was not extensive enough to fully understand the connection. It was clear, however, that to do management, you have to do mapping and transformation. Management allows you to rethink and retool. Our group was interested to know what has happened since this article was written. Have they put this into production? What has changed? It appears there is a follow-up article to this one that would be interesting to read.
The article claims that MARC mapping work is representative of the metadata management task as a whole. Choosing metadata standards based on specific project needs is good, and the projects described here demonstrate how to do that. It's easy to imagine a project where you can start with MARC. But what do you do when no MARC already exists? At IU we have experience in many library departments wth projects that re-use existing MARC metadata.
The group identified three possible cases for metadata management for a digital project: have existing MARC, have existing non-MARC, have no existing structured metadata. Are the strategies outlined in this article useful in all three of these cases? We didn't come to a strong conclusion on this issue.
An interesting discussion grew up around the topic of how to deal with legacy (pre-AACR2) MARC records? Institutional memory is likely the best bet, as documentation comparing older practices and current ones is sparse. Politcal boundaries change and places of publication become no longer correct. Some legacy data is easier to deal with, however. An institution could use an authority vendor to update name headings with death dates. Yet certain data elements should be updated over time, but others shouldn’t. The group noted that most metadata work is bibliographic record based and doesn’t do enough with authority records. Making the full authority structure available to the metadata creation staff is sorely needed.
A substantial amount of discussion time was spent on the topic of collection-specific mappings. The benefits of corse are that these get it done, the way you want it. The drawbacks are potentially reduced shareability and interoperability. One has to take the whole scope of the project in mind to make good decisions and worry about what’s really important. Have to keep USER in mind. This is difficult to do, though. We think “the user needs this information” but we should think “how can the user use this system?” One participant noted that we worry too much about the specialized discovery case to the detriment of the generalized one. How much tweaking of metadata mapping is of use? The community seems to swing back and forth over time between the generalized and specialized approaches.
The discussion then turned more theoretical, with thoughts on the changing roles of libraries – specifically, to what degree should we be the intermediary? If the user is on his or her own, should this change the way we provide access to information? We do see a great deal of evidence that libraries have moved to a model where users interact directly with information with no active intermediation from us. The system provides the intermediation that staff once did. We expect better technologies to automatically enrich our records in the future to help with this. For us, participants felt it was more important to get something out than to get it perfect. We need to make a better effort to integrate authority control into non-MARC environments. Automated methods will rely on the authority records a great deal. It therefore follows that we should send less time on bibliographic records and more on authority work. The MARC world is certainly moving in this direction, with professional catalogers doing more high-value activity, leaving the lower-value tasks to machines or lower-level staff. Mapping activities are an example of the higher-value activity, as seen in this article.
This article describes the most common transformation as MARC to simple DC. To make sure information gets into the right DC fields, one need to understand DC. Those doing the mapping must ask - what is the essential information to go in DC? What really identifies rather than just describes? The role of the cataloger would be to oversee the transformation process, to make sure it works correctly. This would need to happen both on the content end and the technical end.
What should relationship of metadata staff to technical staff be? Metadata staff understand both the source and the target data. They would still have to correct things in the output in the end. It certainly helps if the technical staff understand the data as well. Similarly, metadata staff need to have technical skills. For metadata staff, understanding non-standard source data can be a big challenge. The Bradley films are an example of these challenges here at IU. Each set of materials will have different balance of effort spent on it, based on perceived importance and use. Mapping often unearths mistakes in the original metadata. We must get the best bang for our buck by spending more time on the information that’s really important for the users, and leave the rest alone. Effective projects will also need the involvement of collection development staff.
Summary of MDG Session, 11-19-08
Article read: Cundiff, Morgan V. (2004). "An Introduction to the Metadata Encoding and Transmission Standard (METS)." Library Hi Tech 22(1): 52-64.
The session began with a question raised: is allowing arbitrary descriptive and administrative metadata formats inside METS documents a good idea? The obvious advantage is that it makes METS very versatile. But this could also limit its scope – does that make METS only for digitized versions of physical things, excluding born digital material? The group as a whole didn't believe this was an inherent limitation. The ability to add authorized extension schema over time seems to be a good thing, and necessary for the external schema allowance to work.
The flexibility of METS allows it to be used beyond its textual origins – to scores, sound recordings, images, etc. It could potentially be useful beyond libraries, especially to archives and museums. To balance this flexibility, is knowing some sort of structured metadata is being presented enough to ensure a reasonable level of interoperability?
The discussion then turned to the TYPE attribute on <div>, a topic much discussed in the METS community. How does a METS implementer know what values to use? An organization will presumably develop its own practice but the practices won’t be the same across institutions. A clever name for this was suggested: “plantation” metadata – each place can develop their own.
Are there lessons from library cataloging that could help with this problem? Institutions dealing with the same types of material could join together and harmonize practices. METS Profiles provide the means for documenting this, but they don’t really encourage collaboration. Perhaps the expectation is that the metadata marketplace will converge, and those going their own way will lose out some significant benefits, and see it in their best interest to collaborate.
This line of thought led to the question - How did OCLC/LC/the library community get standardized in the first place? Probably because individuals would write up their own rules, then share them. Eventually these rules became shared practice. Maybe this same shift will happen when sharing really becomes a priority. Diverse practices will converge when people really want them to.
A question was then raised about when METS should be used instead of MARC. When is MARC not enough? A participant made the analogy that this was like comparing a plantation to a video arcade. The two are for different purposes, and METS can include descriptive metadata in any format, including MARC. If you want to allow a certain type of searching, for example, a user wants to search for a recording by a certain group, saying METS is better than MARC doesn't make sense. The descriptive metadata schema used within METS is what is going to make the difference in this case, not the use of METS itself. An implementer will still need good descriptive information.
Participants then noted that we had been talking about systems, but we need to talk more about people. Conversations between communities with different practices will help improve interoperability. Can we standardize access points? To do this we would need to develop vocabularies collaboratively between communities, and talk more so that we understand each other’s point of view.
One participant made an extremely astute observation that the structure of METS makes it seem that it wasn't designed to be used directly by people. While metadata specialists often need to look at METS, and plan for what METS produced by an institution should look like, the commenter is correct that for the most part, METS is intended for machine consumption. A developer present noted that we could write an application that does a lot of what METS does without actually storing it in XML/METS – but the benefit of METS is abstracting out one more layer. Coming full circle to the flexibility issue from earlier in the discussion, it was noted that it is difficult to make standard METS tools (including parsers and generators) due to the almost infinite practices that must be accommodated. This led to the thought that perhaps METS could go much farther in being machine-friendly than it already is. That's a scary thought to metadata specialists who work with it!
The session began with a question raised: is allowing arbitrary descriptive and administrative metadata formats inside METS documents a good idea? The obvious advantage is that it makes METS very versatile. But this could also limit its scope – does that make METS only for digitized versions of physical things, excluding born digital material? The group as a whole didn't believe this was an inherent limitation. The ability to add authorized extension schema over time seems to be a good thing, and necessary for the external schema allowance to work.
The flexibility of METS allows it to be used beyond its textual origins – to scores, sound recordings, images, etc. It could potentially be useful beyond libraries, especially to archives and museums. To balance this flexibility, is knowing some sort of structured metadata is being presented enough to ensure a reasonable level of interoperability?
The discussion then turned to the TYPE attribute on <div>, a topic much discussed in the METS community. How does a METS implementer know what values to use? An organization will presumably develop its own practice but the practices won’t be the same across institutions. A clever name for this was suggested: “plantation” metadata – each place can develop their own.
Are there lessons from library cataloging that could help with this problem? Institutions dealing with the same types of material could join together and harmonize practices. METS Profiles provide the means for documenting this, but they don’t really encourage collaboration. Perhaps the expectation is that the metadata marketplace will converge, and those going their own way will lose out some significant benefits, and see it in their best interest to collaborate.
This line of thought led to the question - How did OCLC/LC/the library community get standardized in the first place? Probably because individuals would write up their own rules, then share them. Eventually these rules became shared practice. Maybe this same shift will happen when sharing really becomes a priority. Diverse practices will converge when people really want them to.
A question was then raised about when METS should be used instead of MARC. When is MARC not enough? A participant made the analogy that this was like comparing a plantation to a video arcade. The two are for different purposes, and METS can include descriptive metadata in any format, including MARC. If you want to allow a certain type of searching, for example, a user wants to search for a recording by a certain group, saying METS is better than MARC doesn't make sense. The descriptive metadata schema used within METS is what is going to make the difference in this case, not the use of METS itself. An implementer will still need good descriptive information.
Participants then noted that we had been talking about systems, but we need to talk more about people. Conversations between communities with different practices will help improve interoperability. Can we standardize access points? To do this we would need to develop vocabularies collaboratively between communities, and talk more so that we understand each other’s point of view.
One participant made an extremely astute observation that the structure of METS makes it seem that it wasn't designed to be used directly by people. While metadata specialists often need to look at METS, and plan for what METS produced by an institution should look like, the commenter is correct that for the most part, METS is intended for machine consumption. A developer present noted that we could write an application that does a lot of what METS does without actually storing it in XML/METS – but the benefit of METS is abstracting out one more layer. Coming full circle to the flexibility issue from earlier in the discussion, it was noted that it is difficult to make standard METS tools (including parsers and generators) due to the almost infinite practices that must be accommodated. This led to the thought that perhaps METS could go much farther in being machine-friendly than it already is. That's a scary thought to metadata specialists who work with it!
Wednesday, November 5, 2008
Summary of MDG Session, 10-16-08
Article read: Eklund, Janice. (2007) "Herding Cats: CCO, XML, and the VRA Core." VRA Bulletin 34, no. 1: 45-68.
The Discussion Group began by picking up a theme from the first meeting of the semester: effective use of terminology in writing about metadata. This article did a good job using new terms consistently and frequently, although terms from VRA Core 3 were occasionally applied to a discussion of Core 4. The discussion of consistency then expanded to consistency in metadata itself. Consistency is very useful when one is combining metadata from multiple sources, and content standards like CCO can go a long way towards promoting this consistency.
The mention of CCO sparked a lively conversation about the way the word “standards” is tossed about in metadata circles. Is CCO a standard or not? CCO and VRA Core are not in total agreement, so what does it mean if both are standards we should follow? One can track why the difference exists—CCO has a broader scope, including museums, than VRA Core. CCO is a standard in the way AACR2 is a standard, but not in the way MARC is a standard. AACR2 is learned by practice, and less by reading the book. CCO is still evolving, taking time to learn and implement. It’s more a guide to best practice than AACR2 is. CCO is principle-based like RDA is supposed to be, because it needs to be applicable to many communities.
The next topic of discussion was whether or not VRA Core is really “core.” Its greater coverage for works of art than Dublin Core certainly speaks to it being a domain-specific “core.” The group was less sure if it represented an “exhaustive core.” Tracking VRA Core’s history could be instructive in this analysis – the evolution from Core 2 to Core 3 to Core 4 shows some stabilization, so this could be evidence that they’ve achieved an agreed-upon core. The only really new thing in Core 4 is the collection root element (in addition to work and image).
The linking capability of VRA Core was singled out as an especially effective part of the format, encouraging the use of identifiers to track relationships within text strings. There is not the infrastructure for collaborative development and sharing of authority records in the visual resources community that there is in the library community, so the process of record linking is more manual now in the VR environment than in the library/MARC community. But there is significant progress being made. The community needs to build good systems, and cooperate between institutions. They also need to expand the notion of authority control, to allow for more variety in name references, for example.
Efforts such as CONA (Cultural Objects Name Authority, forthcoming from the Getty) and the Society of Architectural Historians Architectural Visual Resources Network are helping to build the needed infrastructure. More cooperation overall is needed – the VR community and library community are both starting to realize that each of us having our own copies of records isn’t sustainable. Formats like VRA Core can promote fuller record sharing.
Using separate fields for display and indexing was another feature of VRA Core of interest to the discussion group. It was noted that this practice allowed a great deal of flexibility but also required twice as much work. To decide when this is necessary, one must consider how the information will be used—for search or display? in future systems in addition to future ones? how easy will it be to upgrade systems? It’s more important to include both for data elements that represent key features of the work or medium, for example, cultural context.
The discussion group noted that cultural objects cataloging could be a model for library catalogers looking to re-examine which aspects of their work require the attention of cataloging professionals. Cultural objects cataloging places a greater emphasis on analysis than transcription, which is necessary because cultural objects in general don’t explain themselves. Interestingly enough, some visual resources units are “outsourcing” subject indexing to traditional catalogers. Many catalogers on both sides don’t feel competent to do subjective indexing – is something “about death”? It’s much easier to record form/style, what something is rather than what it is about.
The Discussion Group began by picking up a theme from the first meeting of the semester: effective use of terminology in writing about metadata. This article did a good job using new terms consistently and frequently, although terms from VRA Core 3 were occasionally applied to a discussion of Core 4. The discussion of consistency then expanded to consistency in metadata itself. Consistency is very useful when one is combining metadata from multiple sources, and content standards like CCO can go a long way towards promoting this consistency.
The mention of CCO sparked a lively conversation about the way the word “standards” is tossed about in metadata circles. Is CCO a standard or not? CCO and VRA Core are not in total agreement, so what does it mean if both are standards we should follow? One can track why the difference exists—CCO has a broader scope, including museums, than VRA Core. CCO is a standard in the way AACR2 is a standard, but not in the way MARC is a standard. AACR2 is learned by practice, and less by reading the book. CCO is still evolving, taking time to learn and implement. It’s more a guide to best practice than AACR2 is. CCO is principle-based like RDA is supposed to be, because it needs to be applicable to many communities.
The next topic of discussion was whether or not VRA Core is really “core.” Its greater coverage for works of art than Dublin Core certainly speaks to it being a domain-specific “core.” The group was less sure if it represented an “exhaustive core.” Tracking VRA Core’s history could be instructive in this analysis – the evolution from Core 2 to Core 3 to Core 4 shows some stabilization, so this could be evidence that they’ve achieved an agreed-upon core. The only really new thing in Core 4 is the collection root element (in addition to work and image).
The linking capability of VRA Core was singled out as an especially effective part of the format, encouraging the use of identifiers to track relationships within text strings. There is not the infrastructure for collaborative development and sharing of authority records in the visual resources community that there is in the library community, so the process of record linking is more manual now in the VR environment than in the library/MARC community. But there is significant progress being made. The community needs to build good systems, and cooperate between institutions. They also need to expand the notion of authority control, to allow for more variety in name references, for example.
Efforts such as CONA (Cultural Objects Name Authority, forthcoming from the Getty) and the Society of Architectural Historians Architectural Visual Resources Network are helping to build the needed infrastructure. More cooperation overall is needed – the VR community and library community are both starting to realize that each of us having our own copies of records isn’t sustainable. Formats like VRA Core can promote fuller record sharing.
Using separate fields for display and indexing was another feature of VRA Core of interest to the discussion group. It was noted that this practice allowed a great deal of flexibility but also required twice as much work. To decide when this is necessary, one must consider how the information will be used—for search or display? in future systems in addition to future ones? how easy will it be to upgrade systems? It’s more important to include both for data elements that represent key features of the work or medium, for example, cultural context.
The discussion group noted that cultural objects cataloging could be a model for library catalogers looking to re-examine which aspects of their work require the attention of cataloging professionals. Cultural objects cataloging places a greater emphasis on analysis than transcription, which is necessary because cultural objects in general don’t explain themselves. Interestingly enough, some visual resources units are “outsourcing” subject indexing to traditional catalogers. Many catalogers on both sides don’t feel competent to do subjective indexing – is something “about death”? It’s much easier to record form/style, what something is rather than what it is about.
Tuesday, October 7, 2008
Summary of MDG session, 9-30-08
Article discussed: Greenberg, Jane. (2005). "Understanding Metadata and Metadata Schemes." Cataloging & Classification Quarterly 40, no. 3/4: 17-36.
The discussion began with a general question: Does the MODAL framework appear to be a useful way of evaluating metadata schemas? The group in general thought it was, although expressed concern that some of the language in the article was very academic, which sometimes made it difficult for practicing librarians to follow the argument.
Participants appreciated the fact that some metadata schema such as TEI (p. 28 of the article) have as a stated principle the conversion of resources to newer communication formats. This principle is of great benefit, and would be useful for other metadata schemas as well. Data formats will not stay static - our metadata must adapt its format over time to accommodate new ways of communicating.
Some participants noticed a contrast between the design of metadata schemas based on experience and observation and library cataloging rules that are more formalized and change less frequently. This observation led to the question of whether cataloging rules should be more fluid. When the rules do change, the changes are based on experience. From an implementation point of view, it is difficult both for libraries and our users if the rules are constantly changing. Our legacy data is a very real consideration here. So how do we be flexible and adaptable but at the same time consistent and keep up with the legacy data?
The MODAL framework spoke to participants as an analysis tool - helping evaluate the fitness of a given schema for a given purpose. This gets us away from saying a metadata format is "bad" - rather it lets us say that records using the Dublin Core Metadata Element Set are not well-fit to handle FRBRized data, for example.
The article's methodology of bringing in Cutter's objectives as an example of underlying objectives and principles sat well with the discussion group. One participant noted that not many current studies do this. These assumptions can help us focus our efforts. Follow up work could to do some comparison of Cutter's objectives to different metadata formats.
Terminology issues were a hot topic of discussion at the session. Participants thought some kind of collaboratively-developed metadata glossary would be a good idea. They felt it was important for librarians interested in metadata issues to learn new vocabularies. We need to read more, ingest as much as possible, make connections to what we already do. “Cardinality” was an example of a term which was unfamiliar - it brings in the repeatable vs. not repeatable notion that is familiar, but also covers required/not required. Domains do have specialized vocabularies – they serve as “rites of passage” into various professions. Metadata schemes all have context that assumes a specific knowledge base – this article recognizes that. It would be nice if articles had glossaries, though.
Even with discussion, definitions of some terms did not establish a clear consensus. The term “granularity” was defined in the group as "refinement," "the amount you want to analyze down to,” “extent of the description,” "specificity," and "granular means you can slice in different ways."
Participants appreciated the empirical focus of the article, saying that metadata schema design should be observation/experiment based. It's certainly a good thing to have metadata be practical – actually useful. To help decide what metadata schema to use, try out a couple schemas and see how they work, rather than thinking more abstractly. But also need consider community as a factor. The MODAL framework is “multi-focal” – focusing first on one aspect then go to another. Helps implementers think, for example, about both the community and the data itself.
Participants noted two schools of thought for metadata design: a difference of orientation thinking of a problem looking for a solution, as contrasted with a solution looking for a problem. Is there sill room for cataloger judgment? Absolutely. Perhaps cataloger's judgment is needed more in the application of a content standard rather than a structure standard.
This distinction led participants to speculate whether the line between the two is blurring (although all recognized it has always been somewhat blurry). RDA especially seems to be trying to do both simultneously. One participant noted that libraries seem to be moving to blur the two, while other communities are moving to separate them more.
Is terminology the only barrier to learning more about metadata? Some individuals learn better with theory and others with practice. All need a little of both. It really just takes time – remember what it was like to learn cataloging? Getting out of one's comfort zone is difficult. It’s also difficult to be adventurous, when there is less precedent to follow. It's hard to learn many standards – don’t always know which one to use. When you have to learn lots of things, you learn each of them less well. We also have new objectives, including reaching new people and operating in additional systems. It would be helpful to identify models of other institutions where a technical services unit has made significant progress in these areas.
The group found Table 1, which outlines some typologies of metadata schemas, to be interesting. The lines between them seem arbitrary at worst and murky at best. Over time the thinking in this area has gone from 7 categories to 4 – does this mean our community is looking for simplicity? Does this mean this environment is settling down? Maybe, but initiatives such as the DCMI Abstract Model seem to be going the other direction.
The discussion moved relatively seamlessly from topic to topic, and featured a number of insightful comments, often from new participants. Both nitty-gritty and "big picture" issues were raised. Thanks to all who participated for an enlightening discussion.
The discussion began with a general question: Does the MODAL framework appear to be a useful way of evaluating metadata schemas? The group in general thought it was, although expressed concern that some of the language in the article was very academic, which sometimes made it difficult for practicing librarians to follow the argument.
Participants appreciated the fact that some metadata schema such as TEI (p. 28 of the article) have as a stated principle the conversion of resources to newer communication formats. This principle is of great benefit, and would be useful for other metadata schemas as well. Data formats will not stay static - our metadata must adapt its format over time to accommodate new ways of communicating.
Some participants noticed a contrast between the design of metadata schemas based on experience and observation and library cataloging rules that are more formalized and change less frequently. This observation led to the question of whether cataloging rules should be more fluid. When the rules do change, the changes are based on experience. From an implementation point of view, it is difficult both for libraries and our users if the rules are constantly changing. Our legacy data is a very real consideration here. So how do we be flexible and adaptable but at the same time consistent and keep up with the legacy data?
The MODAL framework spoke to participants as an analysis tool - helping evaluate the fitness of a given schema for a given purpose. This gets us away from saying a metadata format is "bad" - rather it lets us say that records using the Dublin Core Metadata Element Set are not well-fit to handle FRBRized data, for example.
The article's methodology of bringing in Cutter's objectives as an example of underlying objectives and principles sat well with the discussion group. One participant noted that not many current studies do this. These assumptions can help us focus our efforts. Follow up work could to do some comparison of Cutter's objectives to different metadata formats.
Terminology issues were a hot topic of discussion at the session. Participants thought some kind of collaboratively-developed metadata glossary would be a good idea. They felt it was important for librarians interested in metadata issues to learn new vocabularies. We need to read more, ingest as much as possible, make connections to what we already do. “Cardinality” was an example of a term which was unfamiliar - it brings in the repeatable vs. not repeatable notion that is familiar, but also covers required/not required. Domains do have specialized vocabularies – they serve as “rites of passage” into various professions. Metadata schemes all have context that assumes a specific knowledge base – this article recognizes that. It would be nice if articles had glossaries, though.
Even with discussion, definitions of some terms did not establish a clear consensus. The term “granularity” was defined in the group as "refinement," "the amount you want to analyze down to,” “extent of the description,” "specificity," and "granular means you can slice in different ways."
Participants appreciated the empirical focus of the article, saying that metadata schema design should be observation/experiment based. It's certainly a good thing to have metadata be practical – actually useful. To help decide what metadata schema to use, try out a couple schemas and see how they work, rather than thinking more abstractly. But also need consider community as a factor. The MODAL framework is “multi-focal” – focusing first on one aspect then go to another. Helps implementers think, for example, about both the community and the data itself.
Participants noted two schools of thought for metadata design: a difference of orientation thinking of a problem looking for a solution, as contrasted with a solution looking for a problem. Is there sill room for cataloger judgment? Absolutely. Perhaps cataloger's judgment is needed more in the application of a content standard rather than a structure standard.
This distinction led participants to speculate whether the line between the two is blurring (although all recognized it has always been somewhat blurry). RDA especially seems to be trying to do both simultneously. One participant noted that libraries seem to be moving to blur the two, while other communities are moving to separate them more.
Is terminology the only barrier to learning more about metadata? Some individuals learn better with theory and others with practice. All need a little of both. It really just takes time – remember what it was like to learn cataloging? Getting out of one's comfort zone is difficult. It’s also difficult to be adventurous, when there is less precedent to follow. It's hard to learn many standards – don’t always know which one to use. When you have to learn lots of things, you learn each of them less well. We also have new objectives, including reaching new people and operating in additional systems. It would be helpful to identify models of other institutions where a technical services unit has made significant progress in these areas.
The group found Table 1, which outlines some typologies of metadata schemas, to be interesting. The lines between them seem arbitrary at worst and murky at best. Over time the thinking in this area has gone from 7 categories to 4 – does this mean our community is looking for simplicity? Does this mean this environment is settling down? Maybe, but initiatives such as the DCMI Abstract Model seem to be going the other direction.
The discussion moved relatively seamlessly from topic to topic, and featured a number of insightful comments, often from new participants. Both nitty-gritty and "big picture" issues were raised. Thanks to all who participated for an enlightening discussion.
Thursday, June 5, 2008
Summary of MDG session, 5-27-08
Article discussed: Hagedorn, Kat, Suzanne Chapman, and David Newman. (July/August 2007) "Enhancing search and browse using automated clustering of subject metadata." D-Lib Magazine 13, no. 7/8. http://www.dlib.org/dlib/july07/hagedorn/07hagedorn.html
The session began with a brief explanation of the methodology employed by this experiment and the OAI-PMH protocol, as it may not have been clear to those who don’t deal with this sort of technology on a regular basis. After this introduction, discussion moved to wondering why the Michigan “high-level browse list” was chosen for grouping clusters, rather than a more standard list? The group realized the value of a short, extremely general list for this purpose, and noted our own Libraries use a similar locally-developed list. Most standard library controlled vocabularies and classification schemes have far too many top terms to be effective for this sort of use. It was noted that choosing cluster labels, if not the high-level grouping, from a library standard controlled vocabulary would promote interoperability of this enhanced data.
The question of quality control then arose: the article described on person performing a quality check on the cluster labels – this must have been an enormous task! The article mentioned mis-assigned categories that would have been found with a more formal quality review process. Have they thought about how they would fix things on the fly – features like “click here to tell us this is wrong”? Did the experiment designers talk to catalogers or faculty as part of the cluster labeling process? Who were their colleagues they asked to do the labeling?
Is their proposal to not label the clusters at all, but to just connect to the high-level browse categories a good on? The group posited that the high-level browse used the campus structure of majors, rather than not organizational structure of the university. (This is the way the IU Libraries web site is structured). In this case, the subcategories more meaningful than main categories, so at least this level would likely be needed.
The discussion group noted evidence of campus priorities in the high-level browse list, for example that the arts and humanities seemed to be under-represented and lumped together while the sciences received more specific attention. Did this make a difference in the clustering too? As noted in the article, titles in the humanities can be less straightforward than in other discipline, making greater use of metaphors. What do the science records have that humanities records don’t? Abstracts, probably – anything else? Perhaps it’s just that the titles were more specific. Do science subject headings contain more information? Description in humanities collections might be more varied than the language in sciences? Many possibilities were presented but the group wasn’t sure which would really affect the clustering methodology.
The group then wondered if the humanities/sciences differences noted in this article would show up in a single institution, or was it just caused in OAIster because of the fact that different data providers tend to focus on one or the other and the difference is really between data providers rather than between disciplines. The group noted (as a gross generalization) that humanities tend to be more interested in time period, people, and places, whereas the sciences are more interested in topic.
Would the clustering strategy work locally ad not just on aggregations? The suggestion in the article that results might improve if run just on one discipline at a time suggests it might. In this case, clusters would likely be more specific. Perhaps an individual institution could employ this method on full text, and leave running it on metadata records alone to the aggregators. It would be interesting to find out if there’s a difference in effectiveness of this methodology on metadata records for different formats, for example, image vs. text.
The group noted the clustering technique would only be as good as the records from the original site. What if context were missing? (the “on the horse” problem) Garbage in, garbage out, as they say. We understood why the experiment only used English-language records, but it would be interesting to extend this.
The clustering experiment was run using only the data from the title, subject, and description fields. Should they use more? Why not creator? This is useful information. Was it because clusters would then form around creators, which could be collocated using existing creator information? The stopword list was interesting to the group. It made sense why terms such as library and copyright were on it, but there are resources about these things, so we don’t want to artificially exclude them. What if the stopword list were not applied to the title field?
The discussion group wondered how these techniques relate to those operating in the commercial world. Amazon uses “statistically improbable phrases” which seems to be the opposite of this technique – identifying terminology that’s different rather than the same between resources. What about studies comparing these automatic methods to user tagging? No participants knew of such a study in the library literature, but it was noted there might be information on this topic in the information retrieval literature. It would be interesting to compare data from this process to the tags from users generated as part of the LC Flickr project.
The article described the overall approach as attempting to create simple interfaces to complex resources. Is this really our goal? We definitely want to collocate like resources. The interface in the screenshots didn’t seem “Google-style” simple. The group noted that in the library field many believe simple interfaces can only yield simple answers and that people looking with simple techniques are generally just looking for something rather than a comprehensive research goal. This article doesn’t have in its scope a discussion as to whether this is true. One big problem is that the article never defines its user base, ad different user bases employ different search techniques.
The discussion group believed that browseability, as promoted by the clustering technique, is a key idea. With a good browse, the interface can provide more ways to get at resources, and then they are more findable. Hierarchical information can be a good way to get users to resources. With the experiment described in this article, the hierarchy is discipline/genre. Would retrieval improve if we pulled in other data from the record to do faceted browsing? Would this work better for humanities rather than science? Do we need to treat the disciplines differently?
Discussion group participants noted that “this isn’t moonwalking,” meaning that this technique looks promising. It needs some tweaking, but the technique hasn’t promised the moon – it’s not purporting to be a be all, end all solution. Its just something we can do, as one part of the many other techniques we use. Can a simple, Google-style interface eventually work for intensive research needs on this data? Or should it? Should the search just lead them to a seminal article and then they citation chase from there? These are interesting questions.
The group then wondered if the proposal to recluster only every few years was a good one. They would certainly need to do it when getting big new chunks of data that are dissimilar to what’s already in the repository. A possible method would be to randomly test once per month to see if clusters are working out well.
The session ended with some more philosophical questions. Why should services like OAIster exist at all if Google can pick these resources up? Is this type of services beneficial for resources that will never get to the top of a Google search in their native environments? What would happen if one were to apply these techniques to a repository with a more resource-based rather than subject-based collection development policy?
The session began with a brief explanation of the methodology employed by this experiment and the OAI-PMH protocol, as it may not have been clear to those who don’t deal with this sort of technology on a regular basis. After this introduction, discussion moved to wondering why the Michigan “high-level browse list” was chosen for grouping clusters, rather than a more standard list? The group realized the value of a short, extremely general list for this purpose, and noted our own Libraries use a similar locally-developed list. Most standard library controlled vocabularies and classification schemes have far too many top terms to be effective for this sort of use. It was noted that choosing cluster labels, if not the high-level grouping, from a library standard controlled vocabulary would promote interoperability of this enhanced data.
The question of quality control then arose: the article described on person performing a quality check on the cluster labels – this must have been an enormous task! The article mentioned mis-assigned categories that would have been found with a more formal quality review process. Have they thought about how they would fix things on the fly – features like “click here to tell us this is wrong”? Did the experiment designers talk to catalogers or faculty as part of the cluster labeling process? Who were their colleagues they asked to do the labeling?
Is their proposal to not label the clusters at all, but to just connect to the high-level browse categories a good on? The group posited that the high-level browse used the campus structure of majors, rather than not organizational structure of the university. (This is the way the IU Libraries web site is structured). In this case, the subcategories more meaningful than main categories, so at least this level would likely be needed.
The discussion group noted evidence of campus priorities in the high-level browse list, for example that the arts and humanities seemed to be under-represented and lumped together while the sciences received more specific attention. Did this make a difference in the clustering too? As noted in the article, titles in the humanities can be less straightforward than in other discipline, making greater use of metaphors. What do the science records have that humanities records don’t? Abstracts, probably – anything else? Perhaps it’s just that the titles were more specific. Do science subject headings contain more information? Description in humanities collections might be more varied than the language in sciences? Many possibilities were presented but the group wasn’t sure which would really affect the clustering methodology.
The group then wondered if the humanities/sciences differences noted in this article would show up in a single institution, or was it just caused in OAIster because of the fact that different data providers tend to focus on one or the other and the difference is really between data providers rather than between disciplines. The group noted (as a gross generalization) that humanities tend to be more interested in time period, people, and places, whereas the sciences are more interested in topic.
Would the clustering strategy work locally ad not just on aggregations? The suggestion in the article that results might improve if run just on one discipline at a time suggests it might. In this case, clusters would likely be more specific. Perhaps an individual institution could employ this method on full text, and leave running it on metadata records alone to the aggregators. It would be interesting to find out if there’s a difference in effectiveness of this methodology on metadata records for different formats, for example, image vs. text.
The group noted the clustering technique would only be as good as the records from the original site. What if context were missing? (the “on the horse” problem) Garbage in, garbage out, as they say. We understood why the experiment only used English-language records, but it would be interesting to extend this.
The clustering experiment was run using only the data from the title, subject, and description fields. Should they use more? Why not creator? This is useful information. Was it because clusters would then form around creators, which could be collocated using existing creator information? The stopword list was interesting to the group. It made sense why terms such as library and copyright were on it, but there are resources about these things, so we don’t want to artificially exclude them. What if the stopword list were not applied to the title field?
The discussion group wondered how these techniques relate to those operating in the commercial world. Amazon uses “statistically improbable phrases” which seems to be the opposite of this technique – identifying terminology that’s different rather than the same between resources. What about studies comparing these automatic methods to user tagging? No participants knew of such a study in the library literature, but it was noted there might be information on this topic in the information retrieval literature. It would be interesting to compare data from this process to the tags from users generated as part of the LC Flickr project.
The article described the overall approach as attempting to create simple interfaces to complex resources. Is this really our goal? We definitely want to collocate like resources. The interface in the screenshots didn’t seem “Google-style” simple. The group noted that in the library field many believe simple interfaces can only yield simple answers and that people looking with simple techniques are generally just looking for something rather than a comprehensive research goal. This article doesn’t have in its scope a discussion as to whether this is true. One big problem is that the article never defines its user base, ad different user bases employ different search techniques.
The discussion group believed that browseability, as promoted by the clustering technique, is a key idea. With a good browse, the interface can provide more ways to get at resources, and then they are more findable. Hierarchical information can be a good way to get users to resources. With the experiment described in this article, the hierarchy is discipline/genre. Would retrieval improve if we pulled in other data from the record to do faceted browsing? Would this work better for humanities rather than science? Do we need to treat the disciplines differently?
Discussion group participants noted that “this isn’t moonwalking,” meaning that this technique looks promising. It needs some tweaking, but the technique hasn’t promised the moon – it’s not purporting to be a be all, end all solution. Its just something we can do, as one part of the many other techniques we use. Can a simple, Google-style interface eventually work for intensive research needs on this data? Or should it? Should the search just lead them to a seminal article and then they citation chase from there? These are interesting questions.
The group then wondered if the proposal to recluster only every few years was a good one. They would certainly need to do it when getting big new chunks of data that are dissimilar to what’s already in the repository. A possible method would be to randomly test once per month to see if clusters are working out well.
The session ended with some more philosophical questions. Why should services like OAIster exist at all if Google can pick these resources up? Is this type of services beneficial for resources that will never get to the top of a Google search in their native environments? What would happen if one were to apply these techniques to a repository with a more resource-based rather than subject-based collection development policy?
Subscribe to:
Posts (Atom)