Showing posts with label UDC_Summary. Show all posts
Showing posts with label UDC_Summary. Show all posts

Sunday, 26 August 2012

UDC as linked data

The Multilingual UDC Summary has been available as SKOS (XML/RDF) since November 2011. UDC Summary has over 2,500 UDC subdivisions including common auxiliaries (language, place, form, materials, properties etc.). UDC records in this selection contain notes, examples, references and are available in 48 languages. 


In addition to the static export download, the UDC SKOS export is also available for browsing via an html interface which displays each UDC class as a single page with an option for language selection. Both the linked data browsing interface and a single download are available from the UDC linked data webpage. The mapping between UDC and SKOS classes is also available from the same page.

The complete UDC are planned to be made available for machine-to-machine access on the web following the 2012 UDC update. This is planned to include not only 70,000 valid UDC numbers but also 11,000 cancelled classes - which will enable linking and redirecting of library catalogues containing deprecated notations. 

The UDC Summary is now being updated and expanded with UDC MRF 2011 data to include more records in the place auxiliaries and biology areas.

Tuesday, 4 January 2011

Multilingual UDC Summary update - 40 languages to date

UDC Summary (udcS) translations continue with great intensity, using our online, web-based tool and multilingual spreadsheets for offline work. We now have 40 languages online in over 10 different scripts: Armenian, Basque, Bengali, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Esperanto, Estonian, Finnish, French, Galician, Georgian, German, Greek, Hindi, Hungarian, Indonesian, Italian, Japanese, Latvian, Lithuanian, Malayalam, Marathi, Norwegian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swedish, Tamil, Turkish, and Ukrainian.

We expect the top classes in Irish, Punjabi and Vietnamese shortly and we are reaching out to other colleagues interested in collaboration. There are over 100 volunteers, librarians, library school lecturers and researchers currently working on translations in their national language. Database statistics showing the translation progress can be viewed at the translation statistics page. Acknowledgments to contributors are currently noted in a translation tracking table.

There seems to be great interest in using the online schedules even in this present form. Average access to the website in March was over 22,000 hits per day.

Because of the intensity of the translation activity, and in order not to lose momentum, we devoted all of 2010 to encouraging, acquiring, uploading and supporting translations, improving the online translation interface and export tools. You can view a more detailed report in the Extensions & Correction, 31(2009). We are now in the process of designing our web content management system (Drupal) to support management and access to downloading exports, mappings and various useful content we will attach to udcS. Our plan for providing data dumps for download, various kinds of exports, and for publishing udcS as linked data is scheduled for 2011. We have created an alphabetical index of around 10,000 terms in English and we are now looking at creating an interface to this. Both mappings to other systems and natural language access will be our focus in the following months.

Report on the ongoing translation in Italian was provided by Chiara Zara at the ISKO Italy meeting in Venice (1 April 2011) "Il progetto di traduzione multilingue online della CDU".

Sunday, 28 March 2010

Mapping intricacies: UDC to DDC

Last week, I received an email from Yulia Skora (Ukraine) who was interested in the availability of the mapping between UDC Summary and the Summary of the Russian universal classification LBC (BBK - Библиотечно-библиографическая классификация in English: Library Bibliographic Classification) Summary. It reminded me of yet another challenging area of work. When responding to Yulia I realised that the issues with mapping, for instance, UDC Summary to Dewey Summaries [pdf] are often made more difficult because we have to deal with classification summaries in both systems and we cannot use a known exactMatch in many situations.

In 2008, following advice received from colleagues in the HILT project, two of our colleagues quickly mapped 1000 classes of Dewey Summaries to UDC Master Reference File as a whole. This appeared to be relatively simple. The mapping in this case is simply an answer to a question "and how would you say e.g. Art metal work in UDC?"

But when in 2009 we realised that we were going to release 2000 classes of UDC Summary as linked data, we decided to wait until we had our UDC Summary set defined and completed to be able to publish it mapped to the Dewey Summaries.

As we arrived at this stage, little did we realise how much more complex the reversed mapping of UDC Summary to Dewey Summaries would turn out to be.

Mapping the Dewey Summaries to UDC highlighted situations in which the logic and structure of two systems do not agree. Especially because Dewey tends to enumerate combinations of subject and attributes that do not always logically belong together. For instance, 850 Literatures of Italian, Sardinian, Dalmatian, Romanian, Rhaeto-Romanic languages Italian literature. This class mixes languages from three different subgroups of Romance languages. Italian and Sardinian belong to Italo Romance sub-family; Romanian and Dalmatian are Balkan Romance languages and Rhaeto Romance is the third subgroup that includes Friulian Ladin and Romanch. As UDC literature is based on a strict classification of language families, Dewey class 850 has to be mapped to 3 narrower UDC classes 821.131 Literature of Italo-Romance Languages , 821.132 Literature of Rhaeto-Romance languages and 821.135 Literature of Balkan-Romance Languages, or to a broader class 821.13 Literature of Romance languages. Hence we have to be sure that we have all these classes listed in the UDC Summary to be able to express UDC-DDC many-to-one, specific-to-broader relationships.

Another challenge appears when, e.g., mapping Dewey class 890 Literatures of other specific languages and language families, which does not make sense in UDC in which all languages and literatures have equal status. Standard UDC schedules do not have a selection of preferred literatures and other literatures. In principle, UDC does not allow classes entitled 'others' which do not have defined semantic content. If entities are subdivided and there is no provision for an item outside the listed subclasses then this item is subsumed to a top class or a broader class where all unspecified or general members of that class may be expected. If specification is needed this can be divided by adding an alphabetical extension to the broader class. Here we have to find and list in the UDC Summary all literatures that are 'unpreferred' i.e. lumped in the 890 classes and map them again as many-to-one specific-to-broader match.

The example below illustrates another interesting case. Classes Dewey 061 and UDC 06 cover roughly the same semantic field but in the subdivision the Dewey Summaries lists a combination of subject and place and as an enumerative classification, provides ready made numbers for combinations of place that are most common in an average (American?) library. This is a frequent approach in the schemes created with the physical book arrangement, i.e. library shelves, in mind. UDC, designed as an indexing language for information retrieval, keeps subject and place in separate tables and allows for any concept of place such as, e.g. (7) North America to be used in combination with any subject as these may coincide in documents. Thus combinations such as Newspapers in North America, or Organizations in North America would not be offered as ready made combinations. There is no selection of 'preferred' or 'most needed countries' or languages or cultures in the standard UDC edition:



If we map the Dewey Summaries to UDC in general and do not have to worry about a reverse relationship the situation is very simple as shown above.

Mapping of UDC Summary to Dewey Summaries requires more thought.

Firstly, UDC class (7) North America (common auxiliary of place) which simply represents the place has to be mapped to all occurrences in which this place is 'built in' to the Dewey subjects:

063 Organization of North America
073 Journalism of North America
917 Geography of North America
970 History of North America
277 Christianity in North America
317 General Statistics in North America
557 Earth Sciences of North America

The type of mapping from what is a general UDC concept of place (7) North America to a specific subject is clearly a broader-to-narrow match. Mapping of, for instance, UDC class 07 Newspapers. The press (includes journalism) to DDC class of 073 Journalism of North America is again broad-to-narrow match.

Precombined subjects, such as those shown above from Dewey, may be expressed in UDC Summary as examples of combination within various records. To express an exact match UDC class 07 has to contain example of combination 07(7) Journals. The Press - North America. In some cases we have, therefore, added examples to UDC Summary that represent exact match to Dewey Summaries. It is unfortunate that DDC has so many classes on the top level that deal with a selection of countries or languages that are given a preferred status in the scheme, and repeating these preferences in examples of combinations of UDC emulates an unwelcome cultural bias which we have to balance out somehow.

This brings us to another challenge... UDC 913(7) Regional Geography - North America [contains 2 concepts each of which has its URI] is an exact match to Dewey 917 [represented as one concept, 1 URI]. It seems that, because they represent an exact match to Dewey numbers, these UDC examples of combinations may also need a separate URIs so that they can be published as SKOS data.

Albeit challenging, mapping proves to be a very useful exercise and I am looking forward to future work here especially in relation to our plans to map UDC Summary to Colon Classification. We are discussing this project with colleagues from DRTC in Bangalore (India).

UDC Summary - translation in progress for 21 languages

The UDC Summary translation team had a busy week. We have uploaded the top classes for Estonian and Armenian languages, just a day after we uploaded the top classes for Hindi and over 800 classes of Norwegian that we managed to extract from TEKORD data (courtesy of Rurik Greenal).

We now have 21 languages online and over 30 volunteers working on translations.

Our online translation tool is being enhanced as we speak. A browsing list with a colour scheme indicating record completion and enabling easy selection of records for translation was also added last week.

The online editor now allows the editing of a subject index and mapping. Access to this is now available for contributors working in this area.

The translation progress statistics can now be viewed for all 21 languages.

The progress statistics page harvests up-to-the-minute completion statistics for each language from the UDCS database and displays them in graph format using jQuery and jqPlot. The percentage completion figures for each language (compared to English) are shown in tables as the ones exposed on the right.

Saturday, 20 February 2010

UDC Summary: 17 languages online

This weekend we uploaded over 2000 UDC classes in the Ukranian language into the UDC Summary. This is the 17th language so far.

Thanks to help from the publishers and editors of national editions and editors of the UDC Summary, we managed to import almost the complete set of UDC numbers that we needed for many languages.

Our online translator seems to do its job and is being expanded with further features as we speak. Most of the credit, however, goes to our hard working volunteers without whom the whole project would not be possible. We expect that many languages of those that are already online will be completed and proofread by June 2010.

The first alphabetical index and mapping to Dewey summary will appear in March. And we also hope to have the first useful exports available for download. We will be looking for other mappings that may be available and we welcome ideas and suggestions.

Saturday, 19 December 2009

Multilingual UDC Summary - available

The UDC Summary of around 2,000 classes has been online since October 2009 and can now be browsed in 13 languages here (select language in the drop down menu top).

The UDC summary is fully aligned with the UDC MRF 2009 which is going to be released in the following months. This set is made available for free use under the Creative Commons Attribution Share Alike 3.0 license (CC-BY-SA).

BSI, AENOR, CEFAL and VINITI (UDC Consortium members) have supplied their UDC abridged edition data in English, French, Spanish and Russian - which served as a basis for translation work. GFDC "Global Forest Decimal Classification" has given their permission to include their top classes under 630 as an extension to the UDC.

We are adding language data and updates as we speak and changes will be visible on a daily basis. In addition to the above 13 languages, we expect Portuguese and Czech shortly.

Captions in all languages appear first and then scope notes, application notes and example of combinations are added as updates progress.

The following functionalities will be added in the near future: search and ABC relative index browsing (at the moment we edit around 16,000 entries) and a chain index in English will be made available in January.

We will start providing various exports for download as soon as we have three complete and proofread languages.

The effort put into this project by colleagues worldwide is admirable. The entire work put into the UDC Summary so far is entirely voluntary including the programming support, the work of our language editors and translators for which we are most grateful.

As soon as we catch our breath we are going to create our 'hall of fame' page to properly acknowledge all contributors whose names are currently listed on our translation tracking page.

Tuesday, 30 September 2008

On presenting UDC summary on the web

This is in relation with Dan's comment the other day. This made me think that it may be good to make UDC 1000 numbers summary available as exports for various kind of interfaces and experiments. He mentioned Library of Congress Subject Headings presented as linked-date using SKOS vocabulary. I think it was Ed Summers who did this interesting application. Anyway I have hardly been at any vocabulary event this year without someone using this as an example. Alistair Miles showed the view of the full data set at SKOS workshop in July in London - viewing LCSH as a canvas with minuscule dots and zooming into the nodes showed labels and links between concepts.

Subject heading systems do not have hierarchical structure but rather many associative linking hence SKOS made LCSH look as if it makes sense semantically - that is at least at first glance. I am curious how a classification scheme (UDC or any other) would look presented graphically in this fashion in comparison. Each class in classification contains all its subordinated classes and is contained in the class above. In this kind of graphical representation it appears that hierarchy has to be represented as concentric circles. Classifications are completely structured, they have larger amount of paradigmatic/vertical relationships, smaller amount of associative relationships and smaller amount of syntagmatic relationships than this is the case with subject headings.

But maybe and for fun once we will have UDC SKOS export it may be interesting to test this.

Because a 1000 data is only a tip of 70000 database - the linking will not be of the same density as it would be with the entire system but would be interesting. Having said that last year MagnaView created a UDC viewer - i.e. a tool that can be used for visual representation of UDC data for desktop applications and that one was rather interesting - and would be worth developing this approach.

With respect to UDC summary interface ... we are getting there. This one will follow good old tree expansion approach something between FATKS project, Swedish/Finnish/Spanish edition online and BSI UDC online. It will be also interesting if we will have time to try subject-alphabetical index and chain index such as the one in FAT-HUM. I think we will add language by language gradually.