Million Books Workshop (brief report)

Imperial College London.
Friday, March 14, 2008.

David Smith gave the first paper of the morning on “From Text to Information: Machine Translation”. The discussion included a survey of machine translation techniques (including the automatic discovery of existing translations by language comparison), and some of the value of cross-language searching.

[Please would somebody who did not miss the beginning of the session provide a more complete summary of Smith’s paper?]

Thomas Breuel then spoke on “From Image to Text: OCR and Mass Digitisation” (this would have been the first paper in the day, kicking off the developing thread from image to text to information to meaning, but transport problems caused the sequence of presentations to be altered). Breuel discussed the status of professional OCR packages, which are usually not very trainable and have their accuracy constrained by speed requirements, and explained how the Google-sponsored but Open Source OCRopus package intends to improve on this situation. OCRopus is highly extensible and trainable, but currently geared to the needs of the Google Print project (and so while effective at scanning book pages, may be less so for more generic documents). Currently in alpha-release and incorporating the Tesseract OCR engine, this tool currently has a lower error-rate than other Open Source OCR tools (but not the professional tools, which often contain ad hoc code to deal with special cases). A beta release is set for April 2008, which will demo English, German, and Russian language versions, and release 1.0 is scheduled for Fall 2008. Breuel also briefly discussed the hOCR microformat for describing page layouts in a combination of HTML and CSS3.

David Bamman gave the second in the “From Text to Information” sequence of papers, in which he discussed building a dynamic lexicon using automated syntax recognition, identifying the grammatical contexts of words in a digital text. With a training set of some thousands of words of Greek and Latin tree-banked by hand, auto-syntactic parsing currently achieves an accuracy rate something above 50%. While this is still too high a rate of error to make this automated process useful as an end in itself, to deliver syntactic tagging to language students, for example, it is good for testing against a human-edited lexicon, which provides a degree of control. Usage statistics and comparisons of related words and meanings give a good sense of the likely sense of a word or form in a given context.

David Mimno completed the thread with a presentation on “From Information to Meaning: Machine Learning and Classification Techniques”. He discussed automated classification based on typical and statistical features (usually binary indicators: is this email spam or not? Is this play tragedy or comedy?). Sequences of objects allow for a different kind of processing (for example spell-checking), including named entity recognition. Names need to be identified not only by their form but by their context, and machines do a surprisingly good job at identifying coreference and thus disambiguating between homonyms. A more flexible form of automatic classification is provided by topic modelling, which allows mixed classifications and does not require the definition of labels. Topic modelling is the automatic grouping of topics, keywords, components, relationships by the frequency of clusters of words and references. This modelling mechanism is an effective means for organising a library collection by automated topic clusters, for example, rather than by a one-dimensional and rather arbitrary classmark system. Generating multiple connections between publications might be a more effective and more useful way to organise a citation index for Classical Studies than the outdated project that is l’Année Philologique.

Simon Overell gave a short presentation on his doctoral research into the distribution of location references within different language versions of Wikipedia. Using the tagged location links as disambiguators, and using the language cross-reference tags to compare across the collections, he uses the statistics compiled to analyse bias (in a supposedly Neutral Point-Of-View publication) and provide support for placename disambiguation. Overell’s work is in progress, and he is actively seeking collaborators who might have projects that could use his data.

In the afternoon there were two round-table discussions on the subjects of “Collections” and “Systems and Infrastructure” that I may report on later if my notes turn out to be usable.

Posted in Conferences, Projects | Leave a comment

Signs that social scholarship is catching on in the humanities

By way of Peter Suber’s Open Access News:

Spiro, Lisa. “Signs that social scholarship is catching on in the humanities.” Digital Scholarship in the Humanities, March 11, 2008. https://digitalscholarship.wordpress.com/2008/03/11/signs-that-social-scholarship-is-catching-on-in-the-humanities/.

Spiro asks: “To what extent are humanities researchers practicing ‘social scholarship’ … embracing openness, accessibility and collaboration in producing their work?” By way of a provisional answer, she makes observations about “several [recent] trends that suggest increasing experimentation with collaborative tools and approaches in the humanities:”

  1. Individual commitment by scholars to open access
  2. Development of open access publishing outlets
  3. Availability of tools to support collaboration
  4. Experiments with social peer review
  5. Development of social networks to support open exchanges of knowledge
  6. Support for collaboration by funding agencies
  7. Increased emphasis on “community” as key part of graduate education

She also points to the “growth in blogging” and the proliferation of collaborative bibliographic tools.

Posted in Open Source | Leave a comment

Changing the Center of Gravity

Changing the Center of Gravity: Transforming Classical Studies Through Cyberinfrastructure

https://www.rch.uky.edu/CenterOfGravity/

University of Kentucky, 5 October 2007

This is the full audio record of “Changing the Center of Gravity: Transforming Classical Studies Through Cyberinfrastructure”, a workshop funded by the National Science Foundation, sponsored by the Center for Visualization and Virtual Environments at the University of Kentucky, and organized by the Perseus Digital Library at Tufts University.

1) Introduction (05:13)
– Gregory Crane
(download this presentation as an mp3 file – 4.78 MB)

2) Technology, Collaboration, & Undergraduate Research (26:23)
– Christopher Blackwell and Thomas Martin, respondent Kenny Morrell
(download this presentation as an mp3 file – 24.1 MB)

3) Digital Criticism: Editorial Standards for the Homer Multitext (29:02)
– Casey Dué and Mary Ebbott, respondent Anne Mahoney
(download this presentation as an mp3 file – 26.5 MB)

4) Digital Geography and Classics (20:23)
– Tom Elliot, respondent Bruce Robertson
(download this presentation as an mp3 file – 18.6 MB)

5) Computational Linguistics and Classical Lexicography (39:16)
– David Bamman and Gregory Crane, respondent David Smith
(download this presentation as an mp3 file – 35.9 MB)

6) Citation in Classical Studies (38:34)
– Neel Smith, respondent Hugh Cayless
(download this presentation as an mp3 file – 35.3 MB)

7) Exploring Historical RDF with Heml (24:10)
– Bruce Robertson, respondent Tom Elliot
(download this presentation as an mp3 file – 22.1 MB)

8) Approaches to Large Scale Digitization of Early Printed Books (24:38)
– Jeffrey Rydberg-Cox, respondent Gregory Crane
(download this presentation as an mp3 file – 22.5 MB)

9) Tachypaedia Byzantina: The Suda On Line as Collaborative Encyclopedia (20:45)
– Anne Mahoney, respondent Christopher Blackwell
(download this presentation as an mp3 file – 18.9 MB)

10) Epigraphy in 2017 (19:00)
– Hugh Cayless, Charlotte Roueché, Tom Elliot, and Gabriel Bodard, respondent Bruce Robertson
(download this presentation as an mp3 file – 17.3 MB)

11) Directions for the Future (50:04)
– Ross Scaife et al.
(download this presentation as an mp3 file – 45.8 MB)

12) Summary (01:34)
– Gregory Crane
(download this presentation as an mp3 file – 1.44 MB)

Posted in Events, Publications | Leave a comment

CFP: DRHA 2008: New Communities of Knowledge and Practice

By way of a long string of reposts, originally to AHESSC:

Date: Fri, 29 Feb 2008 17:37:17 -0000
From: Stuart Dunn
To: AHESSC@JISCMAIL.AC.UK

CALL FOR PAPERS AND PERFORMANCES

Forthcoming Conference

DRHA 2008: New Communities of Knowledge and Practice

The DRHA (Digital Resources in the Humanities and Arts) conference is held annually at various academic venues throughout the UK. The conference theme this year is to promote discussion around new collaborative environments, collective knowledge and redefining disciplinary boundaries. The conference, hosted by Cambridge with its fantastic choice of conference venues will take place from Sunday 14th September to Wednesday 17th September.

The aim of the conference is to:

  • Establish a site for mutually creative exchanges of knowledge.
  • Promote discussion around new collaborative environments and collective knowledge.
  • Encourage and celebrate the connections and tensions within the liminal spaces that exist between the Arts and Humanities.
  • Redefine disciplinary boundaries.
  • Create a forum for debate around notions of the ‘solitary’ and the collaborative across the Arts and Humanities.
  • Explore the impact of the Arts and Humanities on ICT: design and narrative structures and visa versa.

There will be a variety of sessions concerned with the above but also with a particular emphasis on interdisciplinary collaboration and theorising around practice. There will also be various installations and performances focussing on the same theme. Keynote talks will be given by our plenary speakers who we are pleased to announce are Sher Doruff, Research Fellow (Art, Research and Theory Lectoraat) and Mentor at the Amsterdam School for the Arts, Alan Liu, Professor of English, University of California Santa Barbara and Sally Jane Norman, Director of the Culture Lab, Newcastle University. In addition to this, there will be various round table discussions together with a panel relating to ‘Second Life’ and a special forum ‘Engaging research and performance through pervasive and locative arts projects’ led by Steve Benford, Professor of Collaborative Computing, University of Nottingham. Also planned is the opportunity for a more immediate and informal presentation of work in our ‘Quickfire’ style events. Whether papers, performance or other, all proposals should reflect the critical engagement at the heart of DRHA.

Visit the website for more information and a link to the proposals website.

The Deadline for submissions will be 30 April 2008 and abstracts should be approximately 1000 words.

Cambridge’s venues range from the traditional to the contemporary all situated within walking distance of central departments, museums and galleries. The conference will be based around Cambridge University’s Sedgwick Site, particularly the West Road concert hall, where delegates will have use of a wide range of facilities including a recital room and a ‘black box’ performance space, to cater for this year’s parallel programming and performances.

Sue Broadhurst DRHA Programme Chair

Dr Sue Broadhurst
Reader in Drama and Technology, Head of Drama, School of Arts
Brunel University
West London, UB8 3PH
UK
Direct Line:+44(0)1895 266588 Extension: 66588
Fax: +44(0)1895 269768
Email: susan.broadhurst@brunel.ac.uk.

Posted in Call for papers | Leave a comment

Rieger, Preservation in the Age of Large-Scale Digitization

CLIR (the Council on Library and Information Resources in DC) have published in PDF the text of a white paper by Oya Rieger titled ‘Preservation in the Age of Large-Scale Digitization‘. She discusses large-scale digitization initiatives such as Google Books, Microsoft Live, and the Open Content Alliance. This is more of a diplomatic/administrative than a technical discussion, with questions of funding, strategy, and policy rearing higher than issues of technology, standards, or protocols, the tension between depth and scale (all of which were questions raised during our Open Source Critical Editions conversations).

The paper ends with thirteen major recommendations, all of which are important and deserve close reading, and the most important of which is the need for collaboration, sharing of resources, and generally working closely with other institutions and projects involved in digitization, archiving, and preservation.

One comment hit especially close to home:

The recent announcement that the Arts and Humanities Research Council and Joint Information Systems Committee (JISC) will cease funding the Arts and Humanities Data Service (AHDS) gives cause for concern about the long-term viability of even government-funded archiving services. Such uncertainties strengthen the case for libraries taking responsibility for preservation—both from archival and access perspectives.

It is actually a difficult question to decide who should be responsible for long-term archiving of digital resources, but I would argue that this is one place where duplication of labour is not a bad thing. The more copies of our cultural artefacts that exist, in different formats, contexts, and versions, the more likely we are to retain some of our civilisation after the next cataclysm. This is not to say that coordination and collaboration are not desiderata, but that we should expect, plan for, and even strive for redundancy on all fronts.

(Thanks to Dan O’Donnell for the link.)

Posted in Events, Open Source, Publications, Standards | 1 Comment

Registration: 3D Scanning Conference at UCL

Kalliopi Vacharopoulou wrote, via the DigitalClassicist list:

I would like to draw to your attention the fact that registration for the 3D Colour Laser Scanning Conference at UCL on the 27th and 28th of March has now opened.

The first day (27th of March) will include a keynote presentation and papers on the themes of General Applications of 3D Scanning in the Museum and Heritage Sector and of 3D Scanning in Conservation.

The second day (28th of March) will offer a keynote presentation and papers on the themes of 3D Scanning in Display (and Exhibition) and Education and Interpretation. A detailed programme with the papers and the names of the speakers can be found in our website.

If you would like to attend the conference, I would kindly request to fill in the registration form which you can find in this link and return it to me as soon as possible.

There is no fee for participating (or attending the conference) (coffee and lunch are provided free of charge). Please note that attendance is offered on a first-come, first-served basis.

Please feel free to circulate the information about the conference to anyone who you think might be interested.

In the meantime, do not hesitate to contact me with any inquiries.

Posted in Conferences, Tools | Leave a comment

International Seminar of Digital Philology: Edinburgh, March 25-27, 2008

Seen on the AHeSSC mailing list:

The e-Science Institute Event Announcement

The e-Science Institute is delighted to host the “The Marriage of Mercury and Philology: Problems and Outcomes in Digital Philology”. The conference welcomes both leading scholars and young researchers working on the problems of textual criticism and editorial scholarship in the electronic medium, as well as students, teachers, librarians, archivists, and computing professionals who are interested in representation, access, exchange, management and conservation of texts.

Organiser: Cinzia Pusceddu
Dates and Time: Tuesday 25th March 09.00 – Thursday 27th March 17.00
Place: e-Science Institute
University of Edinburgh
13-15 South College Street
Edinburgh
EH8 9AA

For registration and more details see https://www.nesc.ac.uk/esi/events/854/.

Continue reading

Posted in Conferences | Leave a comment

Palaeographic Image Markup Tools

Does anyone know of any prior work in the area of image markup tools, to enable scholars to markup letterforms (and their constituent strokes) on images of texts?
There is the UVic Image Markup tool:
and the Edition Production and PresentantionTechnology tool:
Dot Porter did a good roundup of the various work going on in this area here: https://www.digitalhumanities.org/dh2007/abstracts/xhtml.xq?id=250
which also points to Digital Incunabula:
as more simple tools to link images and text.
Is there is anyone out there on Stoa using an image markup tool (other than PhotoShop) to trace letter forms over images of text? Any good tools out there that we should know about?
Posted in General | 4 Comments

Search Pigeon

Spotted by way of Peter Suber’s Open Access News:

Search Pigeon is a collection of Google Co-opTM Custom Search Engines (CSEs) designed to make researching on the web a richer, more rewarding, and more efficient process.

Designed for researchers in the Arts and Humanities, with a decidedly interdisciplinary bent, the objective of Search Pigeon is to provide a tool enabling the productive and trustworthy garnering of scholarly articles through customized searching.

Right now SearchPigeon.org provides CSEs that search hundreds of peer-reviewed and open access online journals, provided they are either English-language journals, or provide a translation of their site into English.

Posted in Open Source, Publications, Tools | Leave a comment

Post-doctoral positions and PhD fellowships

Seen and copied from Humanist:

Post-doctoral positions and PhD fellowships in Text
Classification and Automatic Labelling

The Department of Computer Science at Trinity College Dublin is
looking for applications for ONE Postdoctoral positions and TWO PhD
positions in the areas of text classification and automatic labelling
of text streams.

The positions are part of a large research project “Next Generation
Localisation” involving a consortium of leading Irish Universities
(DCU, TCD, UCD and UL) and Industry Partners, funded by the Science
Foundation Ireland (SFI). The project focuses on Language Technology
and Digital Content Management in Localisation. Localisation is the
industrial-scale adaptation of digital content to domain, culture and
language. Successful candidates will join a team of Postdoctoral
researchers, PhD students and research advisors from academia and
industry. Details of the advertised posts are as follows:

– POSTDOCSF32: Postdoctoral Position in “Text Categorisation”

o Description: The successful candidate will research and develop
algorithms for automatic annotation of localisation metadata, and
multilingual text type and genre classification. Candidates must
have a strong background and research record in machine learning
and data-intensive natural language processing, as well as good
programming skills.

o Starting date: 3rd quarter 2008

o Salary: Approx. 38,000-44,000 Euro per annum depending on
experience and qualifications.

For further details, please contact Dr Saturnino Luz
(luzs@cs.tcd.ie) or Dr Carl Vogel (vogel@cs.tcd.ie). To apply,
please email a CV and contact details for two references by March 1,
2008 to Jean.Maypother@cs.tcd.ie. Please include the job reference
(“POSTDOCILT32”) in the subject line of all email correspondence.

————————

– PHDILT33: PhD Fellowship in “Multilingual Text Type and Genre
Classification”

o Description: Candidates must have a strong interest and some
experience in Computational Linguistics and Machine Learning, and
good programming skills.  Experience with syntactic, semantic
and discourse analysis is desirable.

o Starting date: September 2008

o Stipend: Approx. 16,000 Euro per annum (tax exempt) + University
fees (approx. 5,000 Euro per annum) + equipment allowance and a
generous conference travel allowance.

For further details, please contact Dr Carl Vogel
(vogel@cs.tcd.ie). To apply, please email a CV and contact
details for two references by March 1, 2008 to
Jean.Maypother@cs.tcd.ie. Please include the job reference
(“PHDILT33”) in the subject line of all email correspondence.

————————

– PHDILT32: PhD Fellowship in “Automatic Annotation of
Localisation
Metadata”

o Description: Candidates must have a strong interest and
some
experience in Computational Linguistics or Machine Learning,
and
good programming skills.

o Starting date: September
2008

o Stipend: Approx. 16,000 Euro per annum (tax exempt) + University
fees (approx. 5,000 Euro per annum) + equipment allowance and a
generous conference travel
allowance.

For further details, please contact Dr Saturnino Luz
(luzs@cs.tcd.ie). To apply, please email a CV and contact
details for two references by March 1, 2008 to
Jean.Maypother@cs.tcd.ie. Please include the job reference
(“PHDILT32”) in the subject line of all email correspondence.

While the deadline is March 1, 2008, applications will be
considered until the position is filled.

Posted in General | Leave a comment

Digitizing Early Material Culture (CFP)

Posted for Brent Nelson:

Digitizing Early Material Culture: from Antiquity to Modernity
A Seminar to be held in conjunction with
CaSTA (the Canadian Symposium on Text Analysis) 2008:
New Directions in Text Analysis
A Joint Humanities Computing, Computer Science Seminar and Conference at University of Saskatchewan, Saskatoon, 16-18 October 2008
“Digitizing Early Material Culture: from Antiquity to Modernity” seminar will be held at the University of Saskatchewan in Saskatoon 16 October 2008 and will feature guest speakers:

  • Melissa Terras, Lecturer in Electronic Communications in the School of Library, Archive and Information Studies at University College London
  • Lisa Snyder, Associate Director of the Experiential Technologies Centre, University of California Los Angeles

It will be held in conjunction with CaSTA 2008–“New Directions in Text Analysis,” 17-18 August, featuring guest speakers:

  • David Hoover, Professor of English at New York University (keynote)
  • Hoyt Duggan, Professor Emeritus in English at University of Virginia
  • Geoffrey Rockwell, Associate Professor in Humanities Computing at University of Alberta
  • Cara Leitch, PhD candidate in English at University of Victoria

Continue reading

Posted in Call for papers | Leave a comment

Music in TEI SIG

The TEI community have just set up a Special Interest Group for the encoding of music in XML (disclosure: I am one of the moderators). I forward the announcement below:

A Special Interest Group for music encoding in TEI has been created. The goal of the SIG is to examine the current possibilities for encoding both the physical representation of music and the aural common elements between different notation systems, and to decide on a preliminary recommendation/agenda for music encoding in the TEI, whether directly via adoption of new elements or by importing a recommended namespace from an existing external schema.

The discussion will deal with issues like:

  • Encoding western music notation from all time periods, from ancient through modern.
  • Encoding not only the music notation, but the aural aspects common to different notation systems.
  • Encoding music and text together as well as music on its own.

Everyone interested is welcome to participate to our mailing list:
https://listserv.brown.edu/archives/cgi-bin/wa?A0=TEI-MUSIC-SIG
and to our wiki:
https://www.tei-c.org/wiki/index.php/SIG:Music

It is particularly important, I think, that experts in ancient music are represented in this discussion, since many of the participants in the TEI community are mediaeval or modern manuscript scholars. There may (there surely *will*) be features of ancient music that test the limits of standards designed to encode more modern musical notations.

(Does anyone have any nice musical papyri lying around that we could encode in EpiDoc as a test of this sort of markup?)

Posted in Standards | Leave a comment

CFP: Open Scholarship: Authority, Community and Sustainability in the Age of Web 2.0

By way of JISC-Repositories:

The 12th International Conference on Electronic Publishing (25 to 27 June 2008, Toronto, Canada) has just extended its call for papers to 31 January 2008. Full details below …

Continue reading

Posted in Call for papers | Leave a comment

CFP: Virtual Worlds: Libraries, Education and Museums

Gabriel Bodard just posted a call for papers for a “virtual worlds” conference, to be held in Second Life on 8 March 2008. You can read the full CFP in the Digital Classicist Archive. I find it unfortunate that the conference organizers (Bodard is not one) have chosen to organize and publicize the conference via a facebook group that requires interested parties to log in just to read about the event.

Posted in Call for papers | 1 Comment

Humanities GRID Workshop (30-31 Jan; Imperial College London)

By way of the Digital Classicists List:

Epistemic Networks and GRID + Web 2.0 for Arts and Humanities
30-31 January 2008
Imperial College Internet Centre, Imperial College London

https://www.internetcentre.imperial.ac.uk/events

Data driven Science has emerged as a new model which enables researchers to move from experimental, theoretical and computational distributed networks to a new paradigm for scientific discovery based on large scale GRID networks (NSF/JISC Digital Repositories Workshop, AZ 2007). Hundreds of thousands of new digital objects are placed in digital repositories and on the web everyday, supporting and enabling research processes not only in science, but in medicine, education, culture and government.  It is therefore important to build interoperable infra-structures and web-services that will allow for the exploration, data-mining, semantic integration and experimentation of arts and humanities resources on a large scale.  There is a growing consensus that GRID solutions alone are too heavy, and that coupling it with Web 2.0 allows for the development of a more light-weight service oriented architecture (SOA) that can adapt readily to user needs by using on demand utility computing, such as morphological tools, mash-ups, surf clouds, annotation and automated workflows for composing multiple services.  The goal is not just to have fast access to digital resources in the arts and humanities, but to have the capacity to create new digital resources, interrogate data and form hypotheses about its meaning and wider context.  Clearly what needs to emerge is a mixed-model of GRID + Web 2.0 solutions for the arts and humanities which creates an epistemic network that supports a four step iterative process: (i) retrieval, (ii) contextualisation, (iii) narrative and hypothesis building, and (iv) creating contextualised digital resources in semantically integrated knowledge networks.  What is key here is not just managing new data, but the capacity to share, order, and create knowledge networks from existing resources in a semantically accessible form.

To create epistemic networks in the arts and humanities there are core technologies that must be developed.  The aim of this expert METHNET Workshop is to focus on developing a strategy for the implementation of these core technologies on an inter-national scale by bringing together GRID computing specialists with researchers from Classics, Literature and History who have been involved in the creation and use of electronic resources.  The core technologies we will focus on in this two day work-shop are: (i) infrastructure, (ii) named entity, identity and co-reference services, (iii) morphological services and parallel texts, (iv) epistemic networks and virtual research environments.  The idea is to bring together expertise from the UK, US, and European funded projects to agree upon a common strategy for the development of core infra-structure and web-services for the arts and humanities that will enable the use of GRID technologies for advanced research.

DAY ONE- 10:00 – 6:00

SESSION I: GRID + Web 2.0 Infrastructure

SESSION II: Computational and Semantic Services: Named Entity, Identity and Co-reference

  • Paul Watry: Named Entity and Identity Services for the National Archives www.liv.ac.uk
  • Greg Crane –  Co-Reference (Perseus)
  • Hamish Cunningham/Kalina Bontcheva: AKT and GATE: GRID-WEB Services AKT/GATE
  • Martin Doerr – Co-Reference and Semantic Services for Grid + Web 2.0 (FORTH)

DAY TWO: 10:00 – 6:00

SESSION I:  Morphological, Parallel Texts and Citation Services

  • Greg Crane – “Latin Depedency Treebank”, Perseus Project
  • Marco Passarotti – “Index Thomisticus” Treebank
  • Notis Toufexis – ‘Neither Ancient, nor Modern:  Challenges for the creation of a Digital Infrastructure for Medieval Greek’
  • Rob Iliffe – Intelligent Tools for Humanities Researchers, The Newton Project

SESSION II: Epistemic Networks and Virtual Research Environments

Registration fee is £60 and places are limited.

Please contact Dolores Iorizzo (d.iorizzo@ic.ac.uk) to secure a place or for further information.  Please send registration to Glynn Cunin (g.cunin@imperial.ac.uk).

The Imperial College Internet Centre would like to acknowledge generous support from the AHRC METHNET for co-hosting this conference.

Posted in Conferences | Leave a comment