Showing posts with label discoverability. Show all posts
Showing posts with label discoverability. Show all posts

Wednesday, 20 June 2018

Artificial Intelligence: What It Is, How It Works, and What Publishers Can Do with It


Atypon logo
AI was one of the hot topics at last year's ALPSP conference, in this guest blog Hong Zhou, Senior Product Manager for Information Discovery and AI at Atypon give us the 101 on this transformational development.

Artificial intelligence, or AI, is much more than the latest technology buzzword. According to Gartner, by 2020, AI will positively change the behavior of billions of workers and users. And Tata estimates that the vast majority of those workers will work outside of IT.

But what exactly is AI?

AI is a broad set of technologies that use the computational capabilities of machines to “think” like humans. There are many different types of AI, each of which can be used to solve different problems.

So how can AI be employed by scholarly publishers? Ultimately, any publishing technology should make the research experience more productive, increase content usage, and add value to the publisher’s content. To do that, R&D at Atypon explores ways to help readers discover useful and relevant information more quickly by improving search mechanisms and refining content recommendations.

Making content relevant: Recommender systems

Recommender systems will be familiar to anyone who has received suggestions about what other products to buy before or after making an online purchase. Publishers can use them to target relevant products to individual customers by understanding their online site behavior and interests.

Anticipating what readers want: Personalized search

AI-driven recommendation technology can be extended to personalize search as well: reading histories can be used to adjust search rankings specifically to each user—and even suggest new queries that may be relevant—with the goal of understanding a user’s intentions even before they search.

Faster, easier content classification: Semantic auto-tagging

Content tagging underlies many important website capabilities, such as automating the creation of topic-specific pages and content bundles, and powering search results and content recommendations. But tagging documents and maintaining tag sets can be a daunting undertaking. Auto-taggers powered by intelligent machine learning algorithms tag articles accurately and even identify which tags may not be assigned correctly. They save curators time by letting them concentrate their efforts only on content that’s assigned low “confidence scores” by the auto-tagger, thus making it easier for publishers to implement and manage taxonomies.

Content enrichment: Natural language processing

Keywords are traditionally extracted or selected manually, but doing it automatically requires a large amount of training data to identify relationships among topics and key phrases. By enabling machines to understand the meaning of content rather than just the individual words, they can extract more valuable information from content. Natural language processing (NLP) automates key phrase extraction and obviates “teaching” the engine about the content first. By extracting key phrases from different sections of the content and ranking them based on their importance, NLP ultimately improves content categorization and, by extension, content discovery.

Beyond tagging and metadata: Knowledge graphs

A knowledge graph charts all of the possible connections among publication-related information like authors, topics, journals, articles, and even external knowledge databases. Based on these connections, algorithms identify and recommend to researchers the most influential entities, trending topics, and even co-authors and reviewers based on their areas of specialization and the subjects about which they’re writing.

Granular discoverability for text and images: Semantic enrichment

Suppose a researcher wants to interpret many figures associated with a single experiment. Editors have to segment them manually using specialized software—problematic when processing a large number of them. Machine learning can be used to extract sub-figures and captions from compound figures and even separate labels from their associated images, enabling each item to be searched and retrieved individually. Such automation not only reduces the cost of segmentation but also extracts and organizes more valuable information so researchers can search for, compare, and recommend images more precisely and easily.

Search the science, not the text

AI is no longer an aspirational conversation about the future—many of the technologies discussed above are all available today and in use by publishers. By using AI to provide better search results for researchers—and enable publishers to target content more effectively—publishers can deepen researchers’ engagement with their websites, increase the value of their content, and further the pursuit of scientific knowledge by surfacing the information they need more quickly and accurately.


Hong Zhou works on Atypon’s next-generation information discovery technologies. Previously, he was the CTO of Digital Fineprint, a startup that leveraged machine learning algorithms for the insurance industry. He also spent a year designing race car games at Eutechnyx. He holds a PhD in 3D modeling with artificial intelligence algorithms from Aberystwyth University and has published widely on computer science.


Atypon is the proud sponsor of our Awards Dinner at the ALPSP Annual Conference which will take place on 12-14 September this year.


Thursday, 22 February 2018

Highlights from the 2018 University Press Redux Conference: A Sponsor's Perspective


Virtusales was delighted to sponsor the 2018 University Press Redux Conference, which was held at The British Library and completely sold out with an atmosphere that reflected it. Filled with lively discussions on policy, open access, disruptive innovation and the opportunities it presents, the importance of publishers providing genuine added value, and even Brexit and Trump.



In the opening keynote, Timothy Wright, CEO at Edinburgh University Press, presented the challenges for 2018 as: monographs, new models in print and distribution, ebooks, content, the skills gaps, open access, institutional support and the financial challenges facing the industry.

Richard Fisher from Yale University Press said that discoverability remains the number one challenge for most university presses and explained how marketing for individual titles has diminished but is still vital. Michael Jubb of Jubb Consulting reiterated this, stating that 50% of university press sales are sold through global retail channels such as Amazon, making it essential for content to be easily discoverable amidst the proliferation in formats, business models and retailers.

As a software supplier operating globally, we were particularly interested in the parallel session on Global, which looked at university presses outside of Europe, UK and USA. Another compelling topic was Digital, which was covered in the plenary session with Allison Belan from Duke University Press and Charles Watkinson from University of Michigan Press, chaired by Nicole Mitchell from University of Washington Press. The session included stimulating debate on the pros and cons of buying vs building systems and platforms. Allison explained that it is of utmost importance for university presses to understand and own their own content, data and business rules, and Duke’s decision to buy in technology expertise and systems allows the press to focus on what they are good at - creating content.

Author engagement and support was another recurring theme throughout the conference with an emphasis put on the need for publishers to add genuine value throughout the supply chain and for authors, contributors and stakeholders to recognise the value added. This was at the heart of Tuesday’s parallel session on Production where delegates heard from Andy Redman from Oxford University Press, Neil Clarke from CPI UK, Bret Freeman from LifeVroom, and Wednesday’s session on Commissioning with Simon Bell from Emerald Group Publishing, Brian Halley from University of Massachusetts Press and Katherine Reeve from Bath Spa University.

On the policy front, the requirement for all long form works to be available as open access in order to be eligible for the 2027 REF was by far the most impacting point mentioned, with Steven Hill from HEFCE suggesting how it could be achieved using different models such as Freemium, author pays, and mission orientated new university presses.

Closing Keynote: Richard Charkin
Finally, Richard Charkin from Bloomsbury Publishing gave an engaging closing keynote on academic publishing. Technology was mentioned in virtually all presentations, in one form or another, and it seems clear that publishers need its support now more than ever. It is essential for platforms and systems to be flexible enough to manage metadata at content, article and chapter level, bring efficiencies to the production process and supply chain, and liberate publishers so that they can focus on creating, curating and enriching content to engage with their consumers.

At Virtusales we are committed to streamlining publishers’ workflows and bring efficiencies to their business processes with innovative software, exceptional service and a collaborative approach. Our recent white paper looks at the current landscape of academic and educational publishing, some of the disruptions that publishers are facing in today’s arena and ways in which they can capitalise on the changing environment.

Virtusales Publishing Solutions is the creator of the Biblio suite of publishing software and works closely with some of the world's leading academic, scholarly, professional and trade publishers including Harvard University Press, Bloomsbury Publishing, Manchester University Press, Pearson Education, Penguin Random House, Hachette and Macmillan Publishers.


Find out more about Virtusales and the Biblio suite or request our white paper, please visit our website: www.virtusales.com
Follow us on Twitter: @virtusales 
LinkedIn: https://www.linkedin.com/company/virtusales-publishing-solutions/

For further information on the University Press Redux conference and to access slides and audio from the event please visit: https://www.alpsp.org/UPRedux

Thursday, 11 July 2013

Charlie Rapple on leveraging authors' expertise and networks to increase discoverability and usage

Charlie Rapple, co-founder of Kudos
Charlie Rapple, Associate Director at TBI Communications, provided a stark reminder of the environment researchers and publishers operate in at our seminar 'It's all about discoverability, stupid!' held on 10 July.

The problem? There are more and more articles which people have less and less time to read. This results in too many articles that are uncited.

A paper by Michael Mabe and Mayur Amin from 2001 projected 3.26% per annum growth in article numbers (Scientometrics, 51:1 (2001) 147-162) which equates to a lot of papers. In the report The Role of the Critical Review Article in Alleviating Information Overload (Annual Reviews 2011) it was found that 81% of early career researchers felt they should read more of the literature than they do, and 25% suggested they would need to read for more than 24 working hours a week to keep up.

If there is less time for reading, then too many articles will go without being cited. A further article - Five-year impact factor data in the Journal Citation Reports - from 2009 by Péter Jacsó in Online Information Review, explores new indicators based on the level of uncitedness of articles in journals.

The fact that reader time is at a premium is an issue for the reader, author, publisher and librarian. There is a filtering challenge. The standard way of communicating what is in a formal scientific paper is through an abstract, but when you read it, all but an expert might struggle to get to grips with the content. As a result they may miss reading the paper.

One such example is a paper co-authored by R Giles Harrison at the University of Reading entitled 'Electrical signature in polar night cloud base variations' (Environmental Research Letters, March 2013). As a student, if you read the abstract, the first paragraph of which is extracted below, you might not be inspired to go to the full text download.

"Layer clouds are globally extensive. Their lower edges are charged negatively by the fair weather atmospheric electricity current flowing vertically through them. Using polar winter surface meteorological data from Sodankylä (Finland) and Halley (Antarctica), we find that when meteorological diurnal variations are weak, an appreciable diurnal cycle, on average, persists in the cloud base heights, detected using a laser ceilometer."

However, if you view the accompanying video posted on YouTube, you immediately see the headline story, put in simple terms, and communicated by a researcher who is passionate about and committed to his work.



This approach was subsequently picked up in the wider press enabling the paper and author to gain much wider coverage raising the impact of the work.

Harnessing the authors' expertise and explaining research for different audiences is an essential part of this approach. Putting the research in context and explaining why is it important while mapping to human interest or hot topics can help raise the profile of the work and widen the readership. If you can boil down what's unique about it while providing some background on who or what was an inspiration, you can convey the passion behind your work and go beyond formal formats. This is likely to lead to research being cited. Currently, it is time-consuming and manual to draw out the "story" and only a minority of research gets this treatment

Capturing lost value
There are a lot of assets already built on authors' expertise including: lay summaries, impact statements, benefit to society statements, novelty statements etc. All this type of information could be enormously helpful to PR teams and researchers, but is currently held in silos. Part of the problem is the limited infrastructure for capturing and putting these assets to work. There is a real opportunity for the author as a more authoritative hub using existing social media networks such as Facebook and Twitter as well as research-oriented sites such as MyExperiment.org, ResearchObject.org, Mendeley and CiteULike.

Rapple cited the work of Melissa Terras from UCL who blogged and tweeted about uploading papers to the repository, initially to relieve the boredom. She told the story behind the research with every blog post, and was intrigued to note spikes in downloads of her papers that corresponded with her use of social media. This led her to observe:

"If you tell people about your research, they look at it." 

This demonstrates the potential to close the loop, break down the silos and reclaim the value of contextual information that is currently lost or fragmented.

Bringing it all together
These three problems of too much information, not enough time; valuable assets left in silos; and valuable networks being under-utilised require a solution that involves better use of metadata and multimedia. Rapple has been working on a new project, currently in beta, called Kudos, that tries to address this.

Whatever the solution, it will need to take into account:
  • Article level metrics (altmetrics and usage marketing)
  • Author services (open access, academic spring, community, advocacy)
  • Discovery (SEO, social media, international reach, metadata, marketing to individuals)
  • Closing the loop (data driven services, integration not duplication, social machines)
  • Intelligent reading (filtering, multimedia, public accessibility)
  • Research evaluation (funding, REF, STAR, ORCID, H-Index etc)

Wednesday, 10 July 2013

The Royal Society of Chemistry's Richard Kidd with some home truths on semantic enrichment

Richard Kidd on semantics
Richard Kidd, Business Development Manager in the Strategic Innovation Group at the Royal Society of Chemistry, outlined semantics for discovery in STM.

He reflected on the challenge of how you have to learn many different ways to research, as highlighted by Russell Burke from Royal Holloway earlier in the day at 'It's all about discoverability, stupid!' seminar.

How can semantic enrichment help support user search behaviour? And how can it help to show under-researched content? You can use classification, index terms, and identify keywords, then build your own classification.

The value of topic modelling
When RSC Advances launched, it was a new journal covering all of chemistry (and published 1800 articles in 2012). They needed to develop a sensible way of navigating all this content so they used topic modelling. Their publishing expertise gave them 12 broad subjects that would be intuitive to users. They worked with Wordle and then sense checked removing seven topics that were nonsense, one topic that was too general, which left over 120 topics that were classified. The remaining topics can be used for gap analysis, finding hot topics and assessing competitor weaknesses.

Steer clear of ontologies (unless you really want to do them)
A heavier level of classification is ontologies which are used in text mining. An ontology is a machine readable account of what there is in a given domain and how the things there relate to other things. Ontologies tend to be very domain specific and therefore context specific. But in the real world, they don't fit easily into hierarchies which can lead to large gaps in coverage and requires considerable efforts to build and maintain. Kidd's advice? Steer clear of them unless you really want to do them.

Build links on meaning
The semantic web build links on meaning, and each concept should have URLs or URIs. You have to use standard approaches so resources can connect. Why take a semantic approach? There are only a few successful projects out there. Why is that? There are so many different formats, structures, vocabulary, concepts and meanings. The structure should be data-oriented (not HTML). The meaning of data should be clear (not XML). Reusable mappings between data are needed (not XSLT). You should avoid extensive schema rewriting (not data warehouses). And data should have standard APIs (not Flickr). And to quote or paraphrase Lee Harland form Open Phacts, you need to 'synergise with many public efforts.'

Where people have been doing a lot of semantic work, it has tended to be to link their own content together (for example, GeoFacets from Elsevier or Nature.com linked data projects).

Another consideration is text mining. This builds on entity identification (from dictionaries, ontologies, taxonomies) to extract recipes and additional data. You insert the semantic mark-up and conduct sentiment analysis and business intelligence. Examples include NaCTeM tools applied to PMC Europe and SciBite.com.

Does a publisher text mining help text miners? In Kidd's opinion, mostly not.

Semantics won't help with content discovery right now
What should you do? For an immediate return there isn't much evidence that semantics will help your content discovery right now. Customers aren't paying for enriched data. However, making discovery metadata widely available as linked data seems sensible. Semantic works tends to be development rather than discovery. Currently, it is used internally for product development and pump priming. A lot of data is being pushed externally via API for linking, while everyone waits for clever ideas to come along.

And in the future? Big issues are likely to focus on authoring tools and data quality.

'I can get it all for free on Google, yeah?' Russell Burke from Royal Holloway University on information literacy

Russell Burke, Information Consultant
At Royal Holloway University, they spend a lot of time coaching and training potential and current students on information literacy. Russell Burke, an Information Consultant there, outlined what they cover and who they train.

At Royal Holloway they provide training tailored to subject department needs. They cover pre-sessional and sixth formers as well as undergraduates from first to third years. They also train postgraduates including taught masters and PhD researchers and graduate skills. They work with staff although there are varying degrees of engagement between departments. Embedded training is used as well as IS and one-to-one sessions.

Practical sessions are delivered that focus on hands-on activities, using the online resources themselves, and use of more interactive and visual tools such as:

  • Prezi
  • Video tutorials (in-house and YouTube)
  • Online VLE training and quizes using moodle
  • Group work
  • exploring the use of games and gaming 

An example information literacy session for first year undergraduates would include: finding items on your reading lists; developing a search strategy; online resources and services - LibrarySearch, Web of Knowlege, MLA International; searching information sources - hints and tips; citation and referencing - RefWorks; accessing online resources off campus; and using other libraries.

They also help students with finding material on reading lists through their online reading list system, the Moodle (VLE), via handouts (for vintage academics!), and via LibrarySearch. Often, students are thrown when it comes to book chapters as they are so used to searching digitally for journals articles. So they also provide an example of reference to a book. Getting them to think about their essay question or research topic is also covered. The focus is on what they need to find out and searching for information - you can't just type in your essay title. Students are encouraged to ask what it is they are looking for, to think around the topic, and to consider what the essential terms they will need to search for.

Guidance is provided on developing a search strategy. Identify keywords that define your research question. Select relevant information sources. Evaluate and modify your searches and select and save results, then locate copies of promising texts.

Analysing and evaluating information is also crucial. You need a critical evaluation of information sources:

  • Origin - where is your information from? Can you get access to the online full text or print material?
  • Content - is it an academic journal? A published book, a newspaper or a public website or a blog?
  • Relevance - read the abstract (summary) so they know it's relevant.

With citation and referencing, students are guided to acknowledge the author of the source. This will enable the item to be traced (by their lecturer, something most students pick up on first). But it also shows evidence of the scope and depth of your research. Then they have to understand the appropriate reference style (layout etc). They cover legal issues such as plagiarism - something that is incredibly important to the institution - and then provide guidance on selecting and saving results and full text.

Fancy a game?

Other issues that are covered relate to RefWorks: bibliographic management software; capturing, saving and organising references; how to access it via the e-reousrces or an A-Z list; and how to access online self-help tutorials. They also teach them about the various different systems and standards, of which there are many!

Burke closed with a useful tip: they use gaming as an icebreaker for students. It works to engage and present information in a more digestible form and helps with newer students.

Lettie Conrad from SAGE Publications: a publisher's perspective on discoverability

SAGE Publication's Lettie Conrad
Lettie Conrad is Online Publishing Manager at SAGE Publications. At the ALPSP seminar 'It's all about discoverability, stupid! How to get your content seen by the right people', she provided an outline of a publisher's discovery channels and asked: who uses it? why does it matter? and how do you monitor?

Conrad define their channels as those that readers use to find SAGE content. Central to this is understanding their preferences, modes and habits.

They use two primary sources of information:

  1. Market research: usability testing and observation; librarian advisory boards; end user focus groups,surveys etc; information-seeking behavour research studies.
  2. Data analysis: COUNTER reports; Google Analytics; Moz (previously SEOMoz); Data Salon.
Open web search (e.g. Wikipedia, Google). 
Everyone uses it. It is simple and user friendly and drives quantity versus quality traffic providing quick search on a new topic. Ultimately, knowledge of user trends is key and it matters because everyone uses it. Here, SEO = ROI as this is your common 'starter' channel. They monitor with Google Analytics, Moz and market research.

Library search
Users are advanced students, faculty; advanced search and browse; tending towards narrow queries or "known searches". They capture advanced readers and provide win-win discovery services such as ERM feeds, LibGuides, widgets and more. They monitor using Google Analytics, COUNTER - cost/use; and usability testing.

Academic / A&I search 
Users include advanced students, faculty, practitioners; power users; use case: deep research, building expertise (Microsoft Academic search, Google Scholar, SciVeerse Scopus, PAIS, Wilson, etc). This channel matters because they reach experts and 'power users'. It helps with branding and profile in the scholarly ecosystem. It is a mainstream academic search: hybrid, emerging technology (e.g. Microsoft Academic and Google Scholar as game changers) which enables them to reach a wider audience. They monitor with qualitative market research metrics and Google Analytics. Usage from this channel is relatively low. Even so, it is a powerful channel as these are influencers: valuable traffic.

Marketing campaigns
Campaigns used by journals marketing focus on faculty, authors, society members, practitioners and librarians. They receive flyers, emails, register for content alerts, attend conferences, conduct social search. This activity is monitored by market research, email usage, Google Analytics, platform feature usage. With library marketing they contact librarians; generate new sales, focus on renewals and news. They are engaging with faculty and authors to build brand, to influence and to engage. They also contact students and readers through social media and networking. (Two examples are Social Science Space and Social Science Bites Podcast). They monitor through market research, COUNTER reports, Google Analytics and Data Salon.

Metrics and research
Conrad counselled that you get what you pay for with Google Analytics. They also use COUNTER for library ROI, but ultimately, data is not enough: know your users!

HighWire Press's Hugh Blackbourn at 'It's all about discoverability, stupid! How to get your content seen by the right people'

Simon Inger & Hugh Blackbourn field questions
Discoverability is an art and a science. How do you find what you are looking for with fidelity? We have to ensure the content can be found quickly and easily. Otherwise, in today's world, with so many articles competing for the readers attention, if you do not focus on discoverability, it's as if you have not published them.  Hugh Blackbourn, Senior Publication Manager at HighWire Press provided a view from their perspective on discoverability. 

For HighWire Press, there are nine elements to consider with discoverability.

1. Search Engine Optimisation
Rather than building technical challenging 'page turning' PDFs make URLs search engine friendly for spidering. Each new site they launch is announced to the major search engines as they deliver most traffic (e.g. Google, Google Scholar, PubMed, Microsoft, Yahoo and ISI, plus other major reference services such as GeoRef for earth sciences). They meet regularly with Google and Google Scholar to have a good relationship (but it helps they are in California). Google needs to know which journals or articles are accessible to which authors. They use subscriber links which enables access links on Google Scholar.

2. Discoverability - web scale and specialised search
This is a broad category and includes customer specialised search engines, which want data to serve particular customers such as library portals. There are also geographically specialised search engines, which serve a region such as China (e.g. Baidu).

3. Visibility
You need visibility and discoverability to work in harmony. The world has moved on and the journal is no longer as important as it was: people read articles now. They have a widget construction kit (WiCK), use RSS for topic collections, related article recommendation services and social media tagging (Facebook, Twitter, Google +).

4. Search within an issue
1. Quick search
2. Advanced search
3. Within an issue (table of contents)

5. Related content link - standard auto-search
Based on topics or subject collections - you can collect from across titles and set topic yourself.

6. Linking and alerting
Working with NCBI/PubMed, HighWire developed the original technology to have links in the reference section. This led to industry development of CrossRef and the DOI system. They exploit NCBI databases by linking from articles to databases in two way linking. There also had toll-free inter-journal linking, which was an odd form of open access before OA existed.

7. Data distribution and repository deposit 
This supports distribution of publisher metadata to more than thirty direct recipients. include ISI, PubMed, CrossRef etc. They also work with LOCKSS.

8. Text mining
HighWire Press support Open Archive Metadata Harvesting Protocol with content structured for use by text mining researchers. With publisher agreement, they regularly respond to data requests. Text and data mining presents additional challenges: Is a separate sub required? Is it counted for COUNTER or Journal Usage Factor? Is an additional license required etc? Prospect - a new service developed by CrossRef - will be available at the end this year aims to solve individual researcher use-case.

9. The changing market
Technology is changing rapidly with new ways of presenting and accessing content. They deliver content to Amazon for Kindle devices and apps. They have many mobile optimised sites and are able to rapidly deliver mini-sites which include topical sites targeted to a specific audience/membership division which incorporate not just information from the publisher's journals, but external content as well. They can help drive traffic, build membership, increase revenue (e.g. via advertising) and provide benefits to authors and editors. They build apps for Apples iOS and for Android phones and tablets and the use of these devices is restoring the notion that readers browse to find relevant content.

Discoverability has many components - seemingly unending process of refinement and adaption to new technical requiring investment of your time and resources. The results are meaningful, fairly immediate and rewarding. Blackbourn closed by reclaiming the model of 'CPD' - commitment to publishing discovery.

Wednesday, 12 September 2012

ALPSP Conference Day 2: Discovering the needle in a haystack

Ann Lawson introduces the panel

Chaired by Ann Lawson from EBSCO, this session is designed to help publishers understand how they can help academics and professionals to navigate quickly and seamlessly to the trustworthy content they need.

Ann's colleague Harry Kaplanian, Director of Discovery Services at EBSCO Publishing, kicked off with an overview of discovery services as well as the features and benefits for the publishers. 

He began by reminding us of the first discovery system is a library catalogue system in the early 1900s. He then went on to outline the pros and cons of subsequent systems. 

Pros: users can search the entire physical collection quickly; tight ILS integration; one place to search. 
Cons: users can only search catalogue; metadata searching only.

In the 1990s, the first electronic databases began to appear. They aren’t part of the physical collection; change often; multiple tools needed for searching content; students and faculty no longer know where to look; and the e-content just keeps on coming...

Federated search
Pros: single search box for all content; currency of content.
Cons: speed; many indexes; multiple ranking algorithms; larger result sets in complete; internet traffic; content provider traffic.

Web scale discovery
Pros: single search box; search all content, single index and complete result sets
single relevance ranking; speed and bandwidth; no local hardware or software to install; eliminates traditional list problems; drive usage and lower traffic.
Cons: not tightly integrated with ILS.

Usage impact
5000 students
09/10 to 10/11 205% increase in usage in text

He classifies content providers as:
  • Primary publishers
  • Aggregators 
  • Subject Index Providers
  • What do they need to do?
  • What should they watch out for?
Primary Provides - journals & books:
Now standard to provide full text and metadata to discovery services for searching. In most cases content is presented by publishers. The user is guided to full text by link resolver. You have to make sure highest quality metadata and content provided and that active updates to databases are provided.

Aggregators - full journal databases
Most don’t have the right to submit the full text so content presented by aggregator or publishers. The user is guided to full text by link resolver. Make sure highest quality metadata and content is provided, that discovery vendor accurately states if aggregator is actively taking part or not, and active updates to databases are provided

Subject index providers
Powerful subject indexing based on controlled vocabularies, no full text, but a big impact on full Discovery. Make sure the Discovery service is capable of properly searching, merging, ranking and securing entitled access. Consider what happens when a customer cancels subject index subscription, renews, adds it, doesn’t have it?
Make sure Discovery vendor accurately states if subject index provider is actively taking part or not - check the vendor’s claims.


Simon Inger provided an overview of key findings from the Survey on Reader Navigation which is due to be published shortly. Read about the project here. 

Main recommendations for publishers are:
  • Publishers need to support all of the discovery channels that their clients (libraries and readers) want to use
  • Publishers need to understand how different reader types discover and access their content so that they can target readers and authors more effectively
  • Potential to expose more sophisticated discovery information to key channels
  • Potential to differentiate through which discovery channels to make subscriber offers.
The summary report will be published for free later this month. The full report will be published at the same time. All supporting publishers receive these items in return for their help. Full results set and analysis framework will be available for a fee shortly afterwards. 


Robert Faber, Director, Discoverability Program at OUP, concluded with an overview of the Introducing the Oxford Index.


Why does discoverability matter to publishers and librarians? Traffic and use are the lifeblood of digital scholarship. use of subscriptions shows the value of the content. Discovery reveals interest and demand for new content. Customer and user behaviour is changing. If you can’t find it, you won’t use it. People are searching for a topic, not book - 80% of traffic to Oxford Journals is direct to article. There are many search systems and rapid evolution. It's about free content outside the paywall for some products - abstracts, keywords, Oxford Journals, Oxford Scholarship.
The MARC program has been improved and expanded. Linking: some in place, mainly in close neighbourhood. Some editorial linking between products. Some partnerships with other publishers and institutions have been set up. New mobiles sites for Oxford Journals and future products are set up. Library discovery services have been developing new partnerships. SEO is at the centre of product evolution.


Discovery happens in Open Web Search, library services, research hubs, through content links, opt-in services, viral awareness and via OUP web features.

What is the Oxford Index? It’s free discovery from OUP: a standardised description of every item of content, in one place. It incorporates external search partners, an Oxford interface - landing pages: quick pathways to full text, web-searchable; and cross searchable - but they recognise that the website might not continue to be very important. It provides a way to create links and relationships across content with meaningful links that add value and traffic. There are overview pages for quick view of topic links embedded in products. This is a free service integrated with existing products.

What does this mean for search rankings and usage?
  • No change to existing SEO or rankings
  • OI is supplemental route to primary full-text content
  • OI gives Google a super-site map across OUP content
  • Highly-trusted network reinforces destination full-text sites
  • Aim is additional traffic, monitored and reported

What does it mean for library integration? Library visibility? Bring in users from general web search. OI can interact directly with library search. OI identifies library’s provision of full content which highlights the benefits of library services. Visibility through other library, A&I, research services? OI metadata routinely supplied to library systems.

The benefits include i) traffic: sustaining and widening sales, ii) consistent methods - for users and systems, and iii) evolving grid of options to connect content.

Faber finished with trends and predictions that included:
  • importance of india, china and the non-western world
  • differences between journals, book and reference content defined by role/task within the research journey
  • sites that can qualify general users will become a larger focal point of discoverability activity e.g Google Books
  • shift from focus on entrance-point to linking: related content, related services
  • scholarly/research communities play bigger role in academic discoverability.