Showing posts with label harry kaplanian. Show all posts
Showing posts with label harry kaplanian. Show all posts

Wednesday, 7 August 2013

Kurt Paulus on ALPSP International Conference. Key changes, challenges and opportunities

This is the second in a series of reflections on the 2012 ALPSP International Conference by Kurt Paulus, former Operations Director at the Institute of Physics, and long time supporter of ALPSP. Our thanks go to Kurt for capturing the sessions. If this whets your appetite, book by Friday 9 August to avoid the late booking fee.

Key changes, challenges and opportunities

Change is all around us yet the constants traditionally underpinning scholarly publishing largely remain in place: quality control, trusted brands, citability, even mainstream business models like the big deal. Perhaps reassured by that, we can afford to be less fearful when confronting the many changes that are coming our way. There is pressure from governments to review copyright legislation. Governments and funding agencies want to see much more open access than we have been willing to entertain until recently. Technical changes, not necessarily designed with publishers’ needs in mind, are racing ahead and we are forced to go with the flow.


We sometimes forget that our customers also face change. Librarians confront technical and resource constraints and increasingly have to demonstrate how they add value. Authors and readers suffer severe time constraints and have to change their workflows to be most effective. Both need our help. Change is now a constant yet publishers’ reactions can be coloured by fear of it, causing them to cling to outdated models and overlook the undoubted opportunities that change brings with it. ‘Disruptive technologies’ can have a positive impact.

Mobile

Mobile technologies and social networking techniques are ubiquitous and all around us. The growth of consumer-focused mobile technologies such as tablets and smart phones did not initially alert us to the benefits that this might bring to scholarly publishing. Researchers are also consumers, however, and increasingly demand access to their information wherever they might be.

“The scholarly article in 140 characters? No!” Plenary 2 title

While publishers may have been slow off the mark, mobile access is increasing and ‘Apps’ are springing up apace.  Matt Rampone of HighWire gave an overview: in the UK in August this year, over 13% of total online use was via mobile – smartphones, tablets and also e-readers. iOS is the dominant operating system with Android coming up fast. PDF is the format of choice for readers. Average time spent on a site is about 50s (a little longer for a selection of Apps). Mobile usage has a doubling time of 12 months! Rampone’s advice:

  • Invest in mobile.
  • Invest in analytics (know your readers).
  • Create good experiences.

It is fairly clear that publishing on mobile platforms does not yet require a rebirth of the scholarly article but is essential if publishers are to keep pace with readers’ changing behaviours and does give the opportunity to organize and present content in new ways. We mustn’t forget, Tom Reding of BBC Worldwide reminded us, that we are in the business of insight, not journals. The carrot for embracing change is that mobile users are more willing to pay for additional services than desktop users.

Mark Ware’s presentation based on a series of case studies carried out for Outsell confirmed that publishing to mobile platforms is increasingly settling in with STM publishers, starting with the medical and healthcare areas but spreading more widely. The publishers involved included BMJ, Elsevier, Nature, OUP, Wiley-Blackwell and others. Beyond the technical and presentational issues are business ones such as authentication of a user who wants to access a website from a mobile device from within their institution or while on the move. The solutions for RSC Mobile and Oxford Journals Mobile are slightly different for example, and no doubt evolving in response to reader reaction.

Discoverability

There is a data deluge in scholarly publishing, suggested Sophia Ananiadou, chairing one session: information overload, but is that the right metaphor? The deluge is a feature of a system where increased research funding leads to increased publishable output and, even with the most rigorous peer review, to  a growing mountain of stored information. How do we retrieve the stuff we need to progress research further?

“The problem is not information overload, nor filter failure; it is discovery deficit.” Cameron Neylon following Clay Shirky

The issue is not a new one but the scale is, according to Harry Kaplanian, who took us back to card catalogues and OPACs before moving on to more contemporary techniques of finding and retrieving information. We have since moved on to web search, standard references, author generated key words, metadata, A&I services and aggregators to make searches more exhaustive and reliable. A lot is being learned and there is an increasing convergence. Publishers need to be prepared to serve a variety of reader types looking for information through a number of different ‘discovery channels’ (Simon Inger), and need to be on top of their user statistics so that they can better serve their readers and also share their user data with librarians to support their case (value added again).

Oxford University Press has its own Director of the Discoverability Programme, Robert Faber, charged with increasing discoverability of all OUP content across subject areas. This includes among others:

  • Free content outside the paywall – abstracts, keywords etc
  • Improved MARC programmes
  • Enhanced linking between products and services
  • New mobile sites for Oxford Journals
  • Optimized search engines
  • New approaches to library discovery services
  • Analytic tools for tracking user behaviour.

At the top of the tree is the Oxford index, a standardized description of every item of content in one place, an evolving Oxford interface and a way to create links and relationships between content elements.

We are beginning to move towards text and data mining, from a single publisher’s output towards the whole growing corpus of accumulated information, i.e. from looking for a single needle in one haystack to finding a collection of compatible needles in a whole field of haystacks. Anita de Waard, Director of Disruptive Technologies at Elsevier set the scene by talking about working with biologists on how scientists read, how computers read and how they might come together to discover relevant and reliable information rather than just isolated research papers. She sees an evolving future of research communication where researchers compile data, add metadata, overlay the whole thing with workflow tools, then create papers from this material in Google.doc accessible to editors and reviewers, all in the ‘cloud’. Publishers? – we provide the tools!

John McNaught from the National Centre for Text Mining illustrated some of the techniques for adding more value to discovery: natural language searching, searching for acronyms, checking for language nuances, looking for associations, all designed to peel away layer upon layer of increasing complexity to turn unstructured text into structured content linked to other knowledge.

“The value of a network is proportional to the square of the number of connected members.” Metcalfe’s law

Networking and the semantic web continue to be buzzwords. Information is dispersed in many different places. To make sense of it we need structure and context, a resource description framework that identifies objects and connects them, and the appropriate vocabulary. Knowledge networks – associated communities that work on specific topics, linking to move on to a different level – are the next level up from networks of individual papers and reports, according to Stefan Decker.

So what’s the problem?

All this is fascinating work going on at the moment in academic or similar institutions. So will all our problems be solved soon? Not in a hurry, according to Cameron Neylon of PLoS, unless publishers change their ways. The majority of publishers behave as broadcasters of information and are still not thinking of networks of people and tools. The tools publishers have provided are not adequate, and often licences prohibit text mining of material to which the reader already has access. This is an opportunity for publishers to produce premium services. The hardware tools – mobile devices – are already available and are powerful.

“Publishers are too focused on controlling access.” Cameron Neylon

The bottom line for all these speakers is open access, that is having metadata and full text freely accessible with licensing arrangements that permit text mining, if the dream of improved discoverability for researchers is to be achieved. If there is a straw in the wind pointing to how publishers policies may develop it is the recent agreement by ALPSP, STM and the Pharma Documentation Ring (PDR) on a new clause for the PDR Model Licence:

“Text and Data Mining (TDM): download, extract and index information from the Publisher’s Content to which the Subscriber has access under this Subscription Agreement. Where required, mount, load and integrate the results on a server used for the Subscriber’s text mining system and evaluate and interpret the TDM Output for access and use by Authorized Users. The Subscriber shall ensure compliance with Publisher's Usage policies, including security and technical access requirements. Text and data mining may be undertaken on either locally loaded Publisher Content or as mutually agreed.”

Similar sentiments, though in the context of mobile delivery, were expressed by Charlie Rapple later in the conference. Getting quite used to trying to shake audiences out of complacency, Charlie claimed that our users are not happy with us: we have not evolved products and services in line with how they have evolved. New ways of creating, evaluating, curating and distributing information are all around us. We need to win our users back or we will go out of business: if we don’t, someone else will take over!

“Our audience and their needs should direct our strategy.” Charlie Rapple

Critically this means starting with the audience rather than content or devices on which content is delivered. Find out what users need; look into and understand their workflows; deconstruct our content and integrate it with these workflows, making it interactive, relevant and ‘friction-free’.

Kurt Paulus, 2012
Book by Friday 9 August to avoid the late booking fee.

Wednesday, 12 September 2012

ALPSP Conference Day 2: Discovering the needle in a haystack

Ann Lawson introduces the panel

Chaired by Ann Lawson from EBSCO, this session is designed to help publishers understand how they can help academics and professionals to navigate quickly and seamlessly to the trustworthy content they need.

Ann's colleague Harry Kaplanian, Director of Discovery Services at EBSCO Publishing, kicked off with an overview of discovery services as well as the features and benefits for the publishers. 

He began by reminding us of the first discovery system is a library catalogue system in the early 1900s. He then went on to outline the pros and cons of subsequent systems. 

Pros: users can search the entire physical collection quickly; tight ILS integration; one place to search. 
Cons: users can only search catalogue; metadata searching only.

In the 1990s, the first electronic databases began to appear. They aren’t part of the physical collection; change often; multiple tools needed for searching content; students and faculty no longer know where to look; and the e-content just keeps on coming...

Federated search
Pros: single search box for all content; currency of content.
Cons: speed; many indexes; multiple ranking algorithms; larger result sets in complete; internet traffic; content provider traffic.

Web scale discovery
Pros: single search box; search all content, single index and complete result sets
single relevance ranking; speed and bandwidth; no local hardware or software to install; eliminates traditional list problems; drive usage and lower traffic.
Cons: not tightly integrated with ILS.

Usage impact
5000 students
09/10 to 10/11 205% increase in usage in text

He classifies content providers as:
  • Primary publishers
  • Aggregators 
  • Subject Index Providers
  • What do they need to do?
  • What should they watch out for?
Primary Provides - journals & books:
Now standard to provide full text and metadata to discovery services for searching. In most cases content is presented by publishers. The user is guided to full text by link resolver. You have to make sure highest quality metadata and content provided and that active updates to databases are provided.

Aggregators - full journal databases
Most don’t have the right to submit the full text so content presented by aggregator or publishers. The user is guided to full text by link resolver. Make sure highest quality metadata and content is provided, that discovery vendor accurately states if aggregator is actively taking part or not, and active updates to databases are provided

Subject index providers
Powerful subject indexing based on controlled vocabularies, no full text, but a big impact on full Discovery. Make sure the Discovery service is capable of properly searching, merging, ranking and securing entitled access. Consider what happens when a customer cancels subject index subscription, renews, adds it, doesn’t have it?
Make sure Discovery vendor accurately states if subject index provider is actively taking part or not - check the vendor’s claims.


Simon Inger provided an overview of key findings from the Survey on Reader Navigation which is due to be published shortly. Read about the project here

Main recommendations for publishers are:
  • Publishers need to support all of the discovery channels that their clients (libraries and readers) want to use
  • Publishers need to understand how different reader types discover and access their content so that they can target readers and authors more effectively
  • Potential to expose more sophisticated discovery information to key channels
  • Potential to differentiate through which discovery channels to make subscriber offers.
The summary report will be published for free later this month. The full report will be published at the same time. All supporting publishers receive these items in return for their help. Full results set and analysis framework will be available for a fee shortly afterwards. 


Robert Faber, Director, Discoverability Program at OUP, concluded with an overview of the Introducing the Oxford Index.


Why does discoverability matter to publishers and librarians? Traffic and use are the lifeblood of digital scholarship. use of subscriptions shows the value of the content. Discovery reveals interest and demand for new content. Customer and user behaviour is changing. If you can’t find it, you won’t use it. People are searching for a topic, not book - 80% of traffic to Oxford Journals is direct to article. There are many search systems and rapid evolution. It's about free content outside the paywall for some products - abstracts, keywords, Oxford Journals, Oxford Scholarship.
The MARC program has been improved and expanded. Linking: some in place, mainly in close neighbourhood. Some editorial linking between products. Some partnerships with other publishers and institutions have been set up. New mobiles sites for Oxford Journals and future products are set up. Library discovery services have been developing new partnerships. SEO is at the centre of product evolution.


Discovery happens in Open Web Search, library services, research hubs, through content links, opt-in services, viral awareness and via OUP web features.

What is the Oxford Index? It’s free discovery from OUP: a standardised description of every item of content, in one place. It incorporates external search partners, an Oxford interface - landing pages: quick pathways to full text, web-searchable; and cross searchable - but they recognise that the website might not continue to be very important. It provides a way to create links and relationships across content with meaningful links that add value and traffic. There are overview pages for quick view of topic links embedded in products. This is a free service integrated with existing products.

What does this mean for search rankings and usage?
  • No change to existing SEO or rankings
  • OI is supplemental route to primary full-text content
  • OI gives Google a super-site map across OUP content
  • Highly-trusted network reinforces destination full-text sites
  • Aim is additional traffic, monitored and reported

What does it mean for library integration? Library visibility? Bring in users from general web search. OI can interact directly with library search. OI identifies library’s provision of full content which highlights the benefits of library services. Visibility through other library, A&I, research services? OI metadata routinely supplied to library systems.

The benefits include i) traffic: sustaining and widening sales, ii) consistent methods - for users and systems, and iii) evolving grid of options to connect content.

Faber finished with trends and predictions that included:
  • importance of india, china and the non-western world
  • differences between journals, book and reference content defined by role/task within the research journey
  • sites that can qualify general users will become a larger focal point of discoverability activity e.g Google Books
  • shift from focus on entrance-point to linking: related content, related services
  • scholarly/research communities play bigger role in academic discoverability.