Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

Wednesday, 9 September 2015

Anurag Acharya, co-creator of Google Scholar asks: What happens when your library is worldwide and all articles are easy to find?

There was a real sense of anticipation in the room as co-creator of Google Scholar Anurag Acharya stepped up to make the first keynote of the 2015 ALPSP Conference.
Acharya harked back to his time at grad school in 1990. Print was the dominant format. Research had to be physically handled. Every library was limited or bound in different ways. There was wide distribution for core collections, each field would have its own small sets of journals, with wide visibility for published articles. But there was narrow distribution for other journals that were found in far fewer libraries leading to limited visibility for published articles.

Browse was the common way to find research: tables of content for newly arrived issues, bibliography sections of papers you read, shelves of the libraries you could walk to. Some libraries had search services that were often based on titles, authors, keywords, included abstracts. There was no full text indexing or no relevance ranking. The most recent came first. but if you couldn't find it, you can't learn from it! In every way you were limited, by shelves, by institution's budget, that which you don't know about.

Fast forward to 2015. Almost all journals worldwide are online. A large fraction of archives are online. Anyone anywhere can browse it all - let your fingers do the walking. Your library is worldwide - online shelves have no ends. Relevance ranking allows all articles to rise - all articles are equally easy to find, new or old, well-known journal or obscure. Full text indexing allows all sections to rise including conclusions and methods.

Anyone anywhere can find it all: all areas, all languages, all time. Your own area or your colleague's, latest research or well-read classics, free to all users. If you can get online you can join the entire global research community. There is so much more that you can actually read from big deal licenses, free archives, preprints to open access journals and articles.
The transformation is fantastic - he could not have dreamt of this 25 years ago as a grad student. And publishers, societies, libraries and search services have together made this possible.

So how has researcher behaviour changed? What do they look for? What do they read? What do they cite? There is a tremendous growth in queries with many many more users and queries per user in all research and geographical areas.

Queries evolve: there has been the most growth in keyword/concept queries e.g. author name queries, known item queries. The average query length has increased to 4-5 words. There are multiple concepts or entities occur often and most queries are unique. Queries are no longer limited just to their own area. Relevance ranking makes exploration easy and broad queries return classics/seminal work. There's a mix of expert and non-expert queries from users with sustained growth in related area queries. The researcher is no longer limited to narrow areas.

What do they read? There has been steady and sustained growth per user as well as in diversity of areas per user compared to the growth in related areas queries. Users read much more shown through the growth in both abstracts and full texts.

There is more full text available than ever before. Iterative scanning is a common mode: do a query, scan. Abstracts that have full text links in the search interface are selected more frequently, even if they don't actually read the full text. PDF remains extremely popular for full text allowed what is important to the researcher to be accessible to them later.

They have undertaken research into what researchers cite and the evolution of citation patterns. The full report is published on the Google Scholar blog.

Anurag concluded by observing if it is useful, researchers will find/read/cite. The spread of attention is widening across the spectrum to non-elite journals (more specific, less known), older articles, regional journals and dissertations. Good ideas can come from anywhere and insights are not limited to the well-funded or to the web-published. The top 10 journals still publish many top papers: 85% in 1995 to 75% in 2013. The elite are as yet still elite, but less so.

Research is inherently a process of filtering and abstracts are a crucial part of the filtering process. Forcing full text on early-stage users is not useful and limiting COUNTER stats to full text misses much of an article's utility to researchers.

He reflected that we are lucky to live in an era of information plenty. Better a glut than a famine.

Wednesday, 10 July 2013

HighWire Press's Hugh Blackbourn at 'It's all about discoverability, stupid! How to get your content seen by the right people'

Simon Inger & Hugh Blackbourn field questions
Discoverability is an art and a science. How do you find what you are looking for with fidelity? We have to ensure the content can be found quickly and easily. Otherwise, in today's world, with so many articles competing for the readers attention, if you do not focus on discoverability, it's as if you have not published them.  Hugh Blackbourn, Senior Publication Manager at HighWire Press provided a view from their perspective on discoverability. 

For HighWire Press, there are nine elements to consider with discoverability.

1. Search Engine Optimisation
Rather than building technical challenging 'page turning' PDFs make URLs search engine friendly for spidering. Each new site they launch is announced to the major search engines as they deliver most traffic (e.g. Google, Google Scholar, PubMed, Microsoft, Yahoo and ISI, plus other major reference services such as GeoRef for earth sciences). They meet regularly with Google and Google Scholar to have a good relationship (but it helps they are in California). Google needs to know which journals or articles are accessible to which authors. They use subscriber links which enables access links on Google Scholar.

2. Discoverability - web scale and specialised search
This is a broad category and includes customer specialised search engines, which want data to serve particular customers such as library portals. There are also geographically specialised search engines, which serve a region such as China (e.g. Baidu).

3. Visibility
You need visibility and discoverability to work in harmony. The world has moved on and the journal is no longer as important as it was: people read articles now. They have a widget construction kit (WiCK), use RSS for topic collections, related article recommendation services and social media tagging (Facebook, Twitter, Google +).

4. Search within an issue
1. Quick search
2. Advanced search
3. Within an issue (table of contents)

5. Related content link - standard auto-search
Based on topics or subject collections - you can collect from across titles and set topic yourself.

6. Linking and alerting
Working with NCBI/PubMed, HighWire developed the original technology to have links in the reference section. This led to industry development of CrossRef and the DOI system. They exploit NCBI databases by linking from articles to databases in two way linking. There also had toll-free inter-journal linking, which was an odd form of open access before OA existed.

7. Data distribution and repository deposit 
This supports distribution of publisher metadata to more than thirty direct recipients. include ISI, PubMed, CrossRef etc. They also work with LOCKSS.

8. Text mining
HighWire Press support Open Archive Metadata Harvesting Protocol with content structured for use by text mining researchers. With publisher agreement, they regularly respond to data requests. Text and data mining presents additional challenges: Is a separate sub required? Is it counted for COUNTER or Journal Usage Factor? Is an additional license required etc? Prospect - a new service developed by CrossRef - will be available at the end this year aims to solve individual researcher use-case.

9. The changing market
Technology is changing rapidly with new ways of presenting and accessing content. They deliver content to Amazon for Kindle devices and apps. They have many mobile optimised sites and are able to rapidly deliver mini-sites which include topical sites targeted to a specific audience/membership division which incorporate not just information from the publisher's journals, but external content as well. They can help drive traffic, build membership, increase revenue (e.g. via advertising) and provide benefits to authors and editors. They build apps for Apples iOS and for Android phones and tablets and the use of these devices is restoring the notion that readers browse to find relevant content.

Discoverability has many components - seemingly unending process of refinement and adaption to new technical requiring investment of your time and resources. The results are meaningful, fairly immediate and rewarding. Blackbourn closed by reclaiming the model of 'CPD' - commitment to publishing discovery.

Sunday, 20 January 2013

To Disappear or Not to Disappear? How to Avoid Dropping Out of Search Services During a Journal Transition

Ensuring ongoing visibility within search and discovery services is an important success metric for any journal or platform transition but requires careful planning. New sites can easily disappear from search indexes if best practice isn’t followed, resulting in a significant drop in traffic and negative PR. 

ALPSP's webinar To Disappear or Not to Disappear? held last week featured insights from Darcy Dapra, Partner Manager for Google Scholar, who is responsible for external relations with the scholarly publishing and library communities. Highlights from the talk were captured by Louise Russell from Publishing Technology.

Darcy Dapra shared practical advice from a Google Scholar perspective, although the general principles can be applied to other search and discovery services.

The webinar focussed mainly on the implications of a platform change but the advice also applies to any journal transition. The biggest and most common issue identified relates to URL changes. Redirection of URLs from the old to the new site is critical to avoiding being dropped from the search index but is certainly not the full story.

So how exactly do publishers’ avoid dropping out of the search indexes when changing platform? Darcy talked through a comprehensive check list of actions and considerations for publishers and vendors when managing a transition. Most are technical points but some relate more to the user experience and site design and therefore need to be factored into the early stages of planning for any new platform – to share a few nuggets of advice and housekeeping tips from the session:

  • Ensure ALL URLs for a journal site redirect from the old location to the new via HTTP 301 (permanent) redirects for at least 12 months. 
  • Avoid redirecting users to an interstitial page i.e. “this journal has moved” while potentially useful for a reader, it is not useful for robots. 
  • Without redirects your old URLs will effectively disappear from the search indices – the black hole effect! 
  • Article landing pages should include an abstract, ideally above the fold so the user can see it. If not abstract is available a readable image of the first page should be provided. 
  • Ensure the crawler has acces to full-text articles (PDF or HTML). 
  • The browse path or hierarchy for the new site should use HTML links rather than Javascript. Alternative measures should be taken if Javascript is used. 
  • Include metatags to ensure your data is structured well. 
  • If in doubt review with your platform vendor and Google Scholar. 

You can find additional useful resources on Google Scholar pages for basic start up information. The Transfer Code of Practice is a new version of the code that is due to be released shortly and will include further information on URL redirects in particular.