Showing posts with label TDM. Show all posts
Showing posts with label TDM. Show all posts

Thursday, 4 February 2016

All Change in Scholarly Communications: How are the Players – Veterans and Newbies – Adapting?

Fiona Murphy reports from #APE2016
Last month, in characteristically bracing January Berlin weather, around 250 intrepid speakers and delegates attended the 11th Academic Publishing in Europe (APE – pronounced “Ahhhpay”) meeting. Keep an eye on Twitter #ape2016 as all of the presentations were recorded and so should become available in the near future.

A number of familiar characters – large publishers, established platform providers, and so forth – whose language seems to have evolved over the past few years – spoke about ‘openness’ and ‘sharing’ rather than preserving business models. Todd Toler of Wiley, for instance, expressed the “publisher’s value proposition” as having shifted from content provision – basically “moving stuff about” to “strengthening knowledge connections”. This feels like a real turning of tides; such players are now actively aiding and abetting our efforts to garner significant knowledge from our scholarly ecosystem.

In point of fact, there was a general theme around intelligence rather than simply the power of data. Barend Mons bemoaned the existence of “a Christmas tree of hyperlinks and the malpractice of supplementary material’”, instead calling for the training of experts to really understand how machine learning and human interrogation of data can be meshed together to form a powerful whole – “Open Science as a Social Machine” (keep an eye on the IDCC programme in Amsterdam later this month, as he’ll be expanding on the topic there). Meanwhile, Emma Green, of Zapnito – a start-up that aids knowledge-based companies to maximise the impact of their associated experts spoke of growing the ‘knowledge economy’ by reducing the noise and chatter, thereby freeing up the collective intelligence.

John Sack of Highwire’s approach was to examine frictions in the workflow. If workflow is ‘a way of getting things done,’ then instances of friction – with the possible exception of a review stage – largely involve the loss of efficiency. Currently most journal workflows are still based on the original print journal format, but with the version of record shifting online, the resulting misalignments between what is desired and what is produced are causing delays, and infringements of established rules (such as copyright). Friction-reducing tools that can support and simplify the generation, finding, and attribution of scholarly outputs are needed. This can be enabled by standards such as e.g. ORCID or ResearcherID for people, and by initiatives such as openRIF/VIVO for connecting people and their roles to their works and activities. This connectivity will surely boost quality, productivity, and the need for improved garnering of knowledge from our research landscape that generally arose as a theme across APE in general. This connectedness, according to Sack, is about a supported conversation amongst collaborators who are enabled by tools that sift, pre-curate and – potentially – publish their scholarly outputs.

Opportunities for new business models are appearing in a number of points in the workflow – Publons acknowledges and badges peer review activities, Overleaf provides templated support to write journal articles, and Elsevier is leveraging the new Mendeley Data service to enable authors to publish their data and link it immediately with journal articles.

At the same time, policy (=funding) is also moving in the same direction. Stephan Kuster, Head of Policy Affairs for Science Europe explained its function and mission. Science Europe is a think tank set up to support and advise EU National Research Funding Councils around on EU R&D policy issues. Open Access is one of nine key priorities, including enabling authors to hold copyright, supporting sustainable archiving, and publication and dissemination are integral part of research process and should be funded as such.

There was a thoughtful debate about Scholarly Communications Networks and whether they add value, which would not have been possible even a few years ago. Fred Dylla, Emeritus Executive Director of the American Institute of Physics, made the salient point that reputation of the journal still needs to be fundamentally challenged for the landscape to be really disrupted. Currently, the people and institutions making the key decisions about funding, tenure and promotion, are still fixated on journal reputations and impact factors. So, despite feeling as though there has been a lot of progress in the last few years, it also seems there’s still a lot to do.

Luckily there are several opportunities coming up to extend and develop our understanding of and strategies for adapting to this changing landscape. As well as the aforementioned IDCC later this month. And look out for the ALPSP Seminar on research data, digital preservation and innovation in March. Standing on the Digits of Giants is co-organised with the Digital Preservation Coalition and is designed to orientate and empower publishers, research managers and researchers to navigate and flourish in the new landscape.

Another key space to continue these discussions is in the context of the Force11 community, which aims to bring together many of the stakeholders needed at the table to effect change: policy makers, funders, researchers, technologists, publishers, informaticists, lawyers, etc. Force16 promises to be an exciting venue where we’ll be pushing scholarly communications into uncharted territory. Hope to see you there too.

Fiona Murphy, February 2016

Now associated with the Maverick Publishing Specialists, Fiona Murphy has held a range of production and editorial roles at Wiley, Oxford University Press, Random House and Bloomsbury Academic. She specializes in emerging scholarly communications (including Open Science and Open Data) and works to raise expertise and activity levels across the wider research and publications communities. Fiona has written and presented extensively on the research landscape, data and publishing. She is Co-Chair of the World Data System—Research Data Alliance Publishing Data Workflows Working Group, an Editorial Board Member of the Data Science Journal and enjoys organizing meetings. orcid.org/0000-0003-1693-1240

This post was written by Fiona Murphy with the support of Melissa Haendel.



Monday, 18 May 2015

High Value Content: Big Data Meets Mega Text

ALPSP recently updated the Text and Data Mining Member Briefing (member login required). As part of the update, Roy Kaufman, Managing Director of New Ventures at Copyright Clearance Center, provided an overview of the potential of TDM, outlined below.

"Big data may be making headlines, but numbers don’t always tell the whole story. Experts estimate that at least 80 percent of all data in any organization—not to mention in the World Wide Web at large— is what’s known as unstructured data. Examples include email, blogs, journals, Power Point presentations, and social media, all of which are primarily made up of text. It’s no surprise, then, that data mining, the computerized process of identifying relationships in huge sets of numbers to uncover new information, is rapidly morphing into text and data mining (TDM), which is creating novel uses for old- fashioned content and bringing new value to it. Why? Text-based resources like news feeds or scientific journals provide crucial information that can guide predictions about whether the stock market will rise or fall, can gauge consumers’ feelings about a particular product or company, or can uncover connections between various protein interactions that lead to the development of a new drug.

For example, a 2010 study at Indiana University in Bloomington found a correlation between the overall mood of the 500 million tweets released on a given day and the trending of the Dow Jones Industrial Average. Specifically, measurements of the collective public mood derived from millions of tweets predicted the rise and fall of the Dow Jones Industrial Average up to a week in advance with an accuracy approaching 90 percent, according to study author Johan Bollen, Ph.D., an associate professor in the School of Informatics and Computing. At the time, Dr. Bollen predicted, with uncanny accuracy, where he felt TDM was going, from the imprecise, quirky world of Facebook and Twitter to high-value content. He said, "We are hopeful to find equal or better improvements for more sophisticated market models that may in fact include other information derived from news sources and a variety of relevant economic indicators."

In other words, structured data alone is not enough, nor is text mined from the wilds of social media. Wall Street and marketers, eager to predict the right moment to hit buy or sell or to launch an ad campaign, have already moved from mining Facebook and Twitter to licensing high-value content, such as raw newsfeeds from Thomson Reuters and the Associated Press, as well as scientific journal articles reformatted in machine- readable XML. In fact, a 2014 study by Seth Grimes of Alta Plana concludes that the text mining market already exceeds 2 billion dollars per year, with a CAGR of at least 25%.

Far from being irrelevant in our digital age, high-value content is about to have its moment, and not just to improve the odds in the financial world or help marketers sell soap. It represents a new revenue stream for publishers and their thousands of scientific journals as well. For example, in 2003, immunologist Marc Weeber and his associates used text mining tools to search for scientific papers on thalidomide and then targeted those papers that contained concepts related to immunology. They ultimately discovered three possible new uses for the banned drug. “Type in thalidomide and you get between 2,000 and 3,000 hits. Type in disease and you get 40,000 hits,” writes Weeber in his report in the Journal of the American Medical Informatics Association. “With automated text mining tools, we only had to read 100-200 abstracts and 20 or 30 full papers to create viable hypotheses that others could follow up on, saving countless steps and years of research.”

The potential of computer-generated, text-driven insight is only increasing. In his 2014 TedX Talk, Charles Stryker, CEO of the Venture Development Center, points out that the average oncologist, after scouring journals the usual way, reading them one by one, might be able to keep track of six or eight similar cancer cases at a time, recalling details that might help him or her go back, re-read one of two of those articles, and determine the best course of care for a patient with an intractable cancer. The data banks of the two major cancer institutes, on the other hand, hold searchable records of cancer cases that can be reviewed in conjunction with 3 billion DNA base pairs and 20,000 genes contained within each. So using that data would mean a vast improvement in the odds of finding clues to help treat a tricky case or target the best clinical trial for someone with a rare disease. This information might otherwise have been difficult, if not impossible, for even the most plugged-in oncologist to find, let alone read, see patterns, or retain the information for a period of time.

Think, then, of the possibilities of improving healthcare outcomes if the best biomedical research were aggregated in just a few, easily accessible repositories. That’s about to happen. My employer, Copyright Clearance Center (CCC), is coming to market with a new service designed to make it easier to mine high-value journal content. Scientific, technical and medical publishers are opting into the program, and CCC will aggregate and license content to users in XML for text mining. Although the service has not yet fully launched, CCC already has publishers representing thousands of journals and millions of articles participating.

Consider the difficulties of researchers, doctors, or pharmaceutical companies wishing to use text mining to see if cancer patients on a certain diabetes drug might have a better outcome than patients not on the drug. They must go to each publisher, negotiate a price for the rights, get a feed of the journals, and convert that feed into a single useable format. If the top 20 companies did this with the top 20 publishers, it would take 400 agreements, 400 feeds, and 400 XML conversions. The effort would be overwhelming.

Instead, envision a world where users can avail themselves of an aggregate of all relevant journals in their field of interest. Instead of 400 agreements and feeds to navigate and instead of 400 documents to convert to XML, there would be maybe 40 agreements: 20 between the publishers and CCC and 20 with users. There would be no need for customers to convert the text. In other words, researchers could get their hands on the high-value information they need to move research and healthcare forward, in less time, with less effort. And that’s only the beginning. As Stryker said about the promise of TDM, “We are in the first inning of a nine-inning game. It’s all coming together at this moment in time.”

ALPSP Members can login to the website to view the Briefing here.

Roy Kaufman is Managing Director of New Ventures at the Copyright Clearance Center. He is responsible for expanding service capabilities as CCC moves into new markets and services. Prior to CCC, Kaufman served as Legal Director, Wiley-Blackwell, John Wiley and Sons, Inc. He is a member of the Bar of the State of New York and a member of, among other things, the Copyright Committee of the International Association of Scientific Technical and Medical Publishers and the UK's Gold Open Access Infrastructure Program. He formerly chaired the legal working group of CrossRef, which he helped to form, and also worked on the launch of ORCID. He has lectured extensively on the subjects of copyright, licensing, new media, artists' rights, and art law. Roy is Editor-in-Chief of ‘Art Law Handbook: From Antiquities to the Internet’ and author of two books on publishing contract law. He is a graduate of Brandeis University and Columbia Law School.


Friday, 12 September 2014

Welcoming the robots

Mark Bide, Chairman at the Publishers Licensing Society chaired the penultimate panel at the ALPSP International Conference on text and data mining (TDM).

Gemma Hersh, Policy Director at Elsevier talked through the Elsevier TDM policy. It has been controversial with calls to change it. Central to their policy is the use of the ScienceDirect API, designed to help preserve the performance of website for everyone else.

One controversy is that a license is Elsevier's way of exerting control. However, they have a global license (which complies with the UK copyright exception and balances with copyright frameworks). Another complaint is around the click through agreement: critics believe it controls what researchers are doing and takes control away from libraries to place liability on researchers. However, it is an automatic process, there is no additional liability, it is aligned with institutional e-amendment, provides guidelines on reuse and can offer one to one support.

Another complaint is that they didn't allow text mining of images. The reason was they did not hold copyright in all the images so they would do it on request. However, they now do it automatically and include terms of use flagging when they need to contact the copyright owner where it doesn't lie with Elsevier.

There were criticisms that they were trying to claim copyright over TDM output. This was inadvertent and they have adjusted the policy to be a little more flexible and take this into account. A final misconception was that the policy was rigid.

In Europe, they have signed a commitment to facilitate TDM for researchers, but their policy is global. They are also a signatory of CrossRef and think the new service is good.

Mark Bide introduces the panel
Lars Juhl Jensen, based at the Novo Nordisk Foundation Center for Protein Research at the University of Copenhagen, provided an academic perspective on TDM. He considers himself a pragmatic text and data miner. The volume of biomedical research that he has to read is huge. Making sense of structured and unstructured data is key. All he wants to do is data mine. It enables him to do things such as associate diseases and identify conditions. Once you've got the data from text mining, you can then bring it together with experimental data, and from other sources.

As a researcher doing text mining, he needs the text. He doesn't want much else. The format doesn't matter too much. If he can get it in a convenient format, great. The licence has to be reasonable.

Andrew Clark, Associate Director Global Information and Competitive Intelligence Services at UCB,  articulated what TDM means and the part it plays in the scientific industry. He recounted the work of the Pharma Documentation Ring (P-D-R). Their aims are to:

  • Promote exchange of experience/networking among members
  • Encourage commercial development of new information services and systems
  • Jointly assess new and existing products and services
  • Provide a forum for the information industry

Gemma Hersh, Lars Juhl Jensen and Andrew Clark
Literature patent analysis, sentiment analysis and drug safety are just a few of the benefits of TDM. One of the challenges is around the unstructured format that the data comes in at. They need several aggregators to make the data mineable. It's not always easy to get the datasets - from small publishers to large ones. It's quiet expensive and labour intensive.

There are high costs for setting up your data mining. There are a lack of technical skills in the organisation.

There are benefits to TDM that include a managed and in some cases auditable processes for protecting IP. It provides added value and potential new revenues streams. Clark closed with a call for industry collaborations and asked everyone to watch this space.