Showing posts with label discovery. Show all posts
Showing posts with label discovery. Show all posts

Wednesday, 9 September 2015

Researching Researchers: Developing Evidence-Based Strategy for Improved Discovery and Access

How do you improve discovery and access to improve researchers, academics and students better? Roger C Schonfeld, Direct of the Library and Scholarly Communications Program at Ithaka S+R, chaired a panel including publisher, librarian and a library supplier at the 2015 ALPSP Conference.

Lettie Conrad, Executive Manager for Online Products at SAGE talked about their research on discoverability and delivery and learning from users to support their work. It's not about the user experience, it's about understanding the researcher experience. SAGE organises their product delivery on personas based on researchers and use case studies.

Conrad observed that whether we like it or not, the majority of search starts with the mainstream web. As a researcher advances in study skills and moves along their academic careers, they start to shift to speciality databases. Library discovery is for known items.


They undertook research into researcher experience through their workflow. Findings on queries included higher use of open web search reported, validating authenticity, browser trends.  Findings on retrieval included 100% manually managed citations, low use of hyperlinked reference, few 'version of record' checks. Many  use citation metrics, but only if they were above the fold and nearby.

They went on to ask what the uptake was for apps and tools and were surprised to hear that they didn't help with citation. It was a pain point. Easy import of citations was important. Being able to personalise their digital library.  What did this all mean for SAGE's strategy? They take the research findings to help shape strategy and ensure content is discoverable. They ensure they have  good usage statistics. their discovery strategy is based on their channels (library, open web, social media, academic, SAGE universe). Metadata is a key part of their strategy in three ways: stewardship, optimization and distribution. In the future, they are focusing on what's beyond search. What about the serendipitous process?

Deirdre Costello, Senior UX Researcher at EBSCO talked about how user expectations are formed on the open web, what users look for to make decisions about library resources, and why we need to think about our search results as one of the most important user experiences we can craft.

They conducted a video diary research programme to gather honest and open feedback from college and university students aged 14-18 years old. The great thing about this approach is that they saw the whole ecosystem as well as the wider range of tools they use to organise their lives. The expectations from these wider tools get ported on to those for college use.

Students have competing demands on their time from learning to do laundry for the first time, to making friends and keeping in touch with family. In addition to this, the changing neurology of minds to skim and scan content, impacts on how students search and interact with research.

Students have used Google for years and trust it, focusing on the top five results as it must be them that screwed up with the wrong search term, right? It's only when a tutor takes time out to explain how to question sources that students start to understand you can't trust everything you find on the web.




Lisa Janicke Hinchliffe is Professor/Coordinator for Strategic Planning/Coordinator for Information Literacy Services and Instruction at the University of Illinois' Library. They have articulated a user-centric framework of principles for library service development.

If you add a default search to your Easy Search query, there's a massive jump in usage. It is a very important piece of real estate for discovery. They use an evidence based and user centric framework in all their work and repeatedly go back to the data.


Their users value seamless, digital delivery. They want coherent discovery pathways. They want things as simple as possible, but NOT simplistic. When they say they want 'everything', it's from THEIR perspective. They have tried and tested a number of search options: transparency, predictability/explainability  and customisability are important.


Changing user behaviours include: the length of queries are growing, known item searches are increasing and there is an increasing use of copy and paste searching.

The user tasks that they aim to support are:

  • locate known item
  • locate known research tool
  • explore topic
  • identify/access library tools/databases for topic
  • identify/access research data and tools
  • identify assistance.
This had led to a range of discovery principles. They required personalisation and customisation with full library discovery for content, services and spaces. They want the fewest steps from discovery to delivery. Everything owned, licensed or provided by the library should be discoverable. They aim to fully develop and deploy fewer tools. They are aiming for a wide scale implementation of adaptive contextual assistance and use consistent language and labelling. Crucially for a state funded institution, they require the greatest discovery delivered at the lowest cost.

Tuesday, 26 February 2013

ASA 2013: Laura Cox on Pulling Together - Information Flow Throughout the Scholarly Supply Chain

Laura Cox with a messy and complex supply chain
Laura Cox, Chief Marketing Officer at Ringgold Inc, talked through the problems of information flow throughout the scholarly supply chain. If only publishers would use the right identifiers with their content, then there is a huge opportunity to improve information, insight and cost efficiencies.

What are the things that go wrong? Records are unconnected through the supply chain. Links fail between entities, between internal systems, and between external systems. Renewals are mishandled. Journal transfers, access and authentication is mishandled. Authors and individuals are not linked to their institution. Open access fees have to be checked manually. Authors are not linked to their research and funders are not linked to the research they fund.

We need to find a path to using standardized data. Identifiers can help. They can provide a proper understanding of customers, whether author, reader or institution. They also provide a simple basis for wider data governance (that is data governance defined as processes, policies, standards, organization, technology required to organize, maintain and access/interpret data) through:

  • ongoing data maintenance 
  • identifiers enforce uniqueness
  • enable ongoing data governance
  • ensure systems work
  • help with cleansing data for future use.

Cox cited research from Mark Ware and Michael Mabe (The STM Report, 2012) for the wider context:

  • Journals are increasing by 3.5% per annum 
  • There is an increase in the number of articles by 3% per annum
  • The number of researchers is increasing by 3% per annum
  • Growth in China is in double digits
  • There is increasing demand for any time, any where access
  • Library budgets are frozen.

There are a number of identifiers available. For people, there is the International Standard Name Identifier (ISNI) which can apply to authors, playwrights, artists - any creator - which is a bridge identifier. The Open Research and Contributor ID (ORCID) links research to their authors. It disambiguates names looking at the different manner in which they can be recorded and can help remove problems with name changes. They can embed IDs into research workflows and the supply chain, can link to altmetrics and to integrate systems.

Institutional identifiers include Ringgold and ISNI, which map institutions and link data together. This ensures you can identify institutional customers so you can give correct content, and it disambiguates institutional naming conventions.

If you put institutional and author IDs together you gain genuine market intelligence:

  • who's working with whom and where
  • impact of research output on a particular institution - the contribution of their faculty
  • subscription sales or lack thereof
  • where reseach funding is concentrated
  • ability to track open access charges (APCs) to fee structure.

Use internal linking in your systems, you can use identifiers to connect:

  • customer master file
  • financial system
  • CRM/sales database
  • authentication system
  • fulfilment
  • usage statistics
  • submissions system
  • author information.

This enables you to access information from multiple systems in one place, reducing time and cost in locating information, and enabling you to use information to make decisions and inform strategy.

A nice and tidy supply chain
External linking using identifiers will enable you to:

  • ensure accuracy of information
  • speed up data transactions
  • reduce queries
  • reduce costs
  • open data up to new uses
  • provide seamless supply chain where data flows from one org to next
  • ensures that authors receive credit for the work they produce
  • provide a good service to the community.

We need a forum to discuss and pull together: to engage with the problems in data transfer, generate an industry wide policy on using identifiers, break down the data silo mentality, and use universal identifiers to enable our systems to communicate with each other accurately on an ongoing basis. This will help serve the author and reader more effectively and strengthen the links in the supply chain.

ASA 2013: Ed Pentz on CrossMark - A New Era for Metadata

Horse burger, anyone?
Ed Pentz, Executive Director at CrossRef, provided an overview of how CrossMark provides information on enhancements and changes to an article - even if it is downloaded as a PDF to your computer.

With a slide showing a horse-shaped burger, Pentz observed no one knew what was happening in the supply chain and ingredients were mis-labelled. As a consumer it's hard to know what's verified. Third party certification such as Fairtrade or the Soil Association mark have arisen to help consumers. This is an important lesson for the scholarly publishing community.

Pentz is not talking about bibliographic metadata. This is about some of the things that are changing in broader descriptive metadata - what are users starting to ask? They are interested in the status of the content. What's been done to this content? And what can I do with this content?

Good quality metadata drives discovery, however, there are problems with metadata and identification. This is a challenge for primary and secondary publishers as the existing bibliographic supply chain hasn't been sorted, new things being added in, and this could potentially lead to big problems.

NISO announced two weeks ago standards for open access metadata and indicators. The detail is still to follow which will include things like: licensing; has an APC been paid?; if so, how much and who pays it? These factors will be particularly important to help identify open access articles in hybrid journals.

There are a number of new measures that have to be captured via the workflow. These include:

The FundRef Workflow
  • CrossRef has launched the FundRef pilot to provide a standard way of reporting funding sources. 
  • Altmetrics allow you to look at what happens after publication, looks at aspects of usage, post-publication peer review, capturing social buzz and getting beyond impact factors. 
  • PLOS has article level metrics - available via APIs.

What about content changes? Historically, the final version of the record has been viewed as something set in stone. We need to get away from this idea because it doesn't recognise the ongoing stewardship publishers have for the content.

Many things happen to the status of content - post-publication - including:

  • errata
  • updates
  • corrigenda
  • enhancements
  • retractions
  • protocol updates.

As we have heard throughout the conference, the number of retractions are on the rise. Pentz referred back to an article in Nature 478 (2011) on the trouble with science publishing retractions. The case is clear: when content changes, readers need to know, but there is no real system to do this.

In a digital world, notification of changes can be done more effectively, and that's what CrossRef is all about. Another challenge is the use of PDF: there is no way of knowing whether the status has changed. When online, the correction is often listed below the fold, even on a Google search. The whole issue of institutional repositories is also a factor.

What is CrossMark? It is a logo that identifies a publisher maintained copy of a piece of content. When you click on the logo it tells you whether there have been updates, is the copy being maintained by the publishers, where is it publisher maintained, what version is it and other important publication record information.

Taking the example of the PDF sitting on a researcher's hard drive, the document has the CrossMark logo. Click on it for an update on whether the PDF version is current. You can then link through to the clarification if it is there. It includes a status tab and publication record tab. The record tab is a flexible area where publishers can add lots of non-bibliographic information that is useful to reader, for example, peer review, copyright and licensing, FundRef data, location of repository, open access standards, etc.

Lots of things can be enabled by this such as Mendeley. Pentz showed a demo of how a plugin for Google might be written that flags CrossMark when you search. It was launched in April 2012 and has been developing slowly with 50,000 CrossMark pilot deposits since launch, with 400+ updates. They are working with 20+ publishers on CrossMark implementation.