Showing posts with label data mining. Show all posts
Showing posts with label data mining. Show all posts

Monday, 9 October 2017

The Emotion of Data – Your Child Is Always Beautiful


In week’s guest blog we hear from Kent Anderson, the CEO of RedLink and RedLink Network, on the emotional pull of well-presented data….. 


An Excel spreadsheet or data table isn’t usually enough to rouse the emotions. Rigid rows and columns crammed with shapes are difficult to bond with and even harder to get worked up over. Trends are concealed in there somewhere, meaning lurks, yet our senses are stymied by how raw data are assembled.

Over the past 18 months, guiding RedLink, a data company with the slogan “See What You’re Missing,” has opened my eyes to the wonderful emotional pull of well-presented data – what we might call the ultrasound of data, when a real emotional connection begins to occur. I’ve attended dozens of sessions in which we reveal to new customers their data in our products, and every time there is a strong emotional response – the “ooh!” and “wow!” – because they are seeing something of great interest clearly for the first time.

Visualization isn’t the only way to create emotional connections for users. There are other techniques, such as gamification, personalization, and connection.


Visualization – Seeing Is Believing

Turn a set of columns and rows into a set of interactive curves or lines or bars, and suddenly meaning leaps out. Making these trends clear is powerful for sales people, business leaders, managers, and purchasers. There is also the ability RedLink has to import data for libraries and publishers, saving them days or weeks of effort, that liberates time to look at the data and think about its implications.

Gamification – It Makes Data Engaging

Games are great ways to make complex subjects approachable and more understandable. We’ve adopted some aspects of gamification in our products, adding Unlocks and clever names and treasure maps to business-specific products that otherwise would be officious and off-putting. These conceptual candies help to sweeten the experience, adding memorable and pleasant dimensions to the user experience while boosting utility.


Personalization – It’s Your Data

Increasingly, personal data are viewed not as commodities but as elements you have a right to manage. The EU has been more proactive on this front than the US, for example, with initiatives like “the right to be forgotten” and data portability. This places new constraints on data companies. Yet, constraints drive design and innovation, so new services like Remarq – which allows users to put a lot of data about their usage of the scientific and scholarly literature in one place  – are on the leading edge of the data personalization trend.


Connection – Relevance Matters to Meaning

Data matter the most when you can immediately do something with them. We focus a lot on making this happen, whether it’s allowing users to only see data for customers they manage, to see trends across disciplines instead of just around products, to view the macro (consortia, bundles, titles from multiple sources) and the micro (individual institutions, individual titles, individual sources), giving quick paths to relevant views is crucial to making data matter. These views connect the user with the data so that decisions can happen quickly and confidently.

Conclusion

As an independent data company, RedLink helps libraries, consortia, publishers, and end-users “see what they’re missing.” By using visualization, gamification, personalization, and connection, data can become powerful, efficient, and even enjoyable sources of information to help publishers, librarians, administrators, researchers, editors, and authors make better decisions.



Redlink is a proud Silver Sponsor of the ALPSP 2017 Annual Conference.

Thursday, 10 September 2015

What does content and behavioural data mean for publishing? Microsoft's Kuansan Wang considers.

The availability of large amounts of content and behavioural data has also instigated new interdisciplinary research activities in the areas of information retrieval, natural language processing, machine learning, behavioural studies, social computing and data mining.

Kuansan Wang, Director of the Internet Service Research Centre at Microsoft Research considered the impact for the publishing and consumption of content, drawing on observations derived from a web scale data set, newly released to the public.

If you think about the web as a gigantic library of the future, then you should think about the semantic web as the librarian. It involves trust, proof, logic, ontology vocabulary, rdf schema, xml schema, Unicode and URI.

A central theme for the semantic web is trying to help a machine read and makes sense: human readable versus machine readable contents. The semantic web requires humans to define a standard for data formats and models. It has an explicit and precise specification of knowledge representation that everyone has to agree upon.

The knowledge web is where a machine reads human readable contents. With the knowledge web, the machine learns to conflate different formats of the same thing. It involves latent and fuzzy representation of knowledge learned by mining big data.

There has been a paradigm shift in discovery. Traditional web search involves index keywords in documents, matches keywords in queries and has the relevance of "10 blue links". With knowledge web search it digests the world's knowledge, matches user intent and has a dialogue experience.

The dialogue acts in Bing and Cortana are:
  1. answer 
  2. confirmation 
  3. disambiguation 
  4. suggestion 
  5. progress: refinement.

In Bing, you get answers, there is an element of confirmation/correction, refinement dialogue and digressive suggestion. The interface is designed for naturally spoken language with context, confirmation and answer. You don't have to go to the search page, the disambiguation starts as you type. They train the system to try to summarise what it has to learn.

Some of the issues that bug the academic community are:
  • How to recommend completions for seldom observed or never foreseen queries?
  • How to rank these suggestions?
  • How to avoid making suggestions leading to no or bad results?
For finding researchers and potential collaborators they train a machine to go through and aggregate all the information.
Cortana provides proactive suggestions on Windows Android IOS. Concept is based on the successful personal assistants to the stars who write down the interests and activities of the people they serve to gain better insight. They have built in a lot of switches you can turn on/off for personalisation and if you have privacy concerns and now trained Cortana to do this for academics. One of the pain points you hit as a researcher is that you hit a paywall. Cortana tries to help by showing not only the academic article, but also related news stories.

The latest Microsoft vision is about empowering every person and every business to achieve more. They intend to do this through re-imaged productivity, more personal computing and most intelligent cloud. This translates to academic search, Cortana Academic and Project Oxford.



Wednesday, 5 August 2015

ALPSP Awards Spotlight on… RightFind™ XML for Mining from Copyright Clearance Center

In this second of the ALPSP Awards for Innovation in Publishing finalists' posts,
Jake Kelleher, Senior Director, Licensing and Business Development at Copyright Clearance Center (CCC) talks about RightFind™ XML for Mining, a solution which facilitates copyright-compliant access to full-text article content for text and data mining.

ALPSP: Tell us a bit about your company

CCC was founded over 35 years ago at the suggestion of the United States Congress to provide an efficient market for the clearing of photocopy rights. Today, CCC is a global leader in content and licensing solutions for publishers, businesses and academic institutions. Located near Boston, we have more than 380 employees dedicated to serving the needs of over 12,000 publishers and providing innovative content and rights-licensing technology solutions for more than 35,000 customers around the globe.

In the mid-90s, when publishers were struggling with how to best protect copyrighted material in a new online world, they worked closely with CCC to create RightsLink®, which made it easy for visitors to a publisher’s website to purchase permissions and other services. Years later, we collaborated with our publisher customers again to use RightsLink’s advanced ecommerce capabilities to manage article processing charges (APCs) and other author fees for journal articles. As a result, we launched our RightsLink for Open Access platform with a range of new capabilities developed with input from customers and partners. This summer, we introduced our new RightFind XML for Mining Service in partnership with six major publishers.

ALPSP: What is the project that you submitted for the Awards?

We submitted our latest advance, RightFind XML for Mining, a great example of CCC’s culture of innovation offering new revenue opportunities for life science publishers.


ALPSP: Tell us more about how it works and the team behind it

XML for Mining is built on the RightFind platform, CCC’s unique suite of cloud-based corporate workflow solutions that offer immediate access to a full range of STM peer-reviewed journal content.

XML for Mining launches from this platform and contains normalized, full-text articles from multiple rightsholders, allowing researchers to create a corpus of articles relevant to their research. Once the corpus has been created, researchers download these articles into their text mining solutions. XML for Mining almost eliminates the weeks, or even months, of work it takes to prepare scientific content for use in text mining solutions, thus accelerating research and discovery.

In fact, we recently received some very positive feedback from a leading international pharmaceutical company that they love the full-text search and can’t get it anywhere else. As always with CCC, feedback and ideas from our customers guide everything we do.   We developed RightFind XML for Mining with input from text mining researchers and publishers looking for a voluntary, market-based licensing solution.  And we will continue to listen to the market, add new publisher content and develop new features, and explore ways to collaborate with technology partners.

ALPSP: Why do you think it demonstrates publishing innovation?

RightFind XML for Mining represents a huge step forward for publishers navigating text and data mining waters. For these publishers, the solution has several key advantages. It enables end users to have access to aggregated article content from multiple rightsholders in a single service with normalized metadata, and that has never been done before. It offers a web-based, user-friendly interface, as well as a RESTful Application Programming Interface (API), so that researchers can easily identify and download, in full-text XML format, relevant articles for mining within their workflow.



It also reduces the necessity for time-consuming one-off licensing negotiations with publishers, along with the associated costs to administer, format and deliver custom content feeds to individual customers.  Moreover, our robust API enables integration with leading text and data mining software platforms.  Thus, publishers gain a new channel for their content along with valuable usage data, while users can reduce the number of operational steps involved with finding a needle in the haystack.

ALPSP: What are your plans for the future?

We create content workflow and licensing solutions that make copyright work for everyone. That means identifying and reducing pain points for publishers and their customers. RightFind XML for Mining reaffirms CCC’s role as an independent intermediary between content providers and users that can identify inefficiencies and create bridges between the two groups.  As publisher and market needs evolve, we will evolve our services, as well.

Jake Kelleher is Senior Director, Licensing and Business Development at Copyright Clearance Center. RightFind™ XML for Mining – Copyright Clearance Center’s text mining solution lets life sciences researchers move to a deeper level of discovery– beyond abstracts to direct access to full-text articles in XML format. XML for Mining saves time usually spent acquiring, licensing, and converting articles from individual publishers. Now, researchers can identify and download article collections from multiple publishers through a single, normalized source that is copyright-compliant.

The winner of the ALPSP Awards for Innovation in Publishing, sponsored by Publishing Technology, will be announced at the ALPSP Conference. Book now to secure your place.

Tuesday, 23 September 2014

Big data: mining or minefield? Kurt Paulus reflects...

Who's Afraid of Big Data? Not this panel...
"Data are the stuff of research: collection, manipulation, analysis, more data… They are also the stuff of contemporary life: surveillance data, customer data, medical data… Some are defined and manageable: a researcher’s base collection from which the paper draws conclusions. Some are big: police data, NHS data, GCHQ data, accumulated published data in particular fields. Two Plenaries and several other papers at ALPSP 2014 addressed the issues, options, opportunities and some threats around them.

There have long been calls for authors’ underlying research data to be made accessible, so as to substantiate research conclusions, suggest further work and so on. The main Plenaries concerned themselves with Big Data, usually unstructured sets of elements of unprecedented scale and scope, such as the whole of Wikipedia, accumulated Google searches, the biomedical literature, the visible galaxies in the universe. The challenge of ‘mining’ these datasets is to bring structure to them so that new insights emerge beyond those arising from limited or sampled data. This requires automation, big computing resources and appropriate but speeded-up human intervention and sometimes crowd sourcing.

Gemma Hersh from Elsevier on TDM
Text and data mining has some kinship with information discovery where usually structured datasets are queried, but goes well beyond it by seeking to add structure to very large sets of data where there is no or little structure, so that information can be clustered, trends identified and concepts linked together to lead to new hypotheses. The intelligence services provide a prime, albeit hidden example. So does the functional characterization of proteins, the mining of the UK Biobank for trends and new insights or the crowd-sourced classification of galaxy types to test cosmological theories.

Inevitably there are barriers and issues. The data themselves are often inadequate; for example not all drug trials are published and negative or non-results are frequently excluded from papers. Research data are not always structured and standardized and authors are often untutored in databases and ontologies. The default policy, it was recommended, should be openness in the availability of authors’ published and underlying data, standardized with full metadata and unique identifiers, to make data usable and mitigate the need for sophisticated mining.

CrossRef's Rachael Lammey
Because of copyright and licensing, not all data are easily accessed for retrieval and mining. Increasingly licensing for ‘non-commercial purposes’ is permitted but exactly what is non-commercial is ill-defined, particularly in pharmaceuticals. Organizations like CrossRef, CCC, PLS and others are beginning to offer services that support the textual, licensing and royalty gathering processes for both research and commercial data mining.

Rejecting the name tag Cassandra, Paul Uhlir of the National Academies urged a note of caution. Big Data is changing the public and academic landscape, harbouring threats of disintermediation, complexity, luddism and inequality and exposing weaknesses in reproducibility, scientific method, data policy, metrics and human resources, amongst others.


Paul F. Uhlir urges caution

Judging by the remainder of these sessions and the audience reaction, excitement was more noticeable than apprehension.

ALPSP of course is on the ball and has just issued a Member Briefing on Text and Data Mining (member login required) and will publish a special online-only issue of Learned Publishing before the end of this month."





Kurt Paulus
kurtpaulus@hotmail.co.uk

Monday, 5 November 2012

Bill Matthews discusses data analytics on 15 November



HighWire is excited about contributing to the ALPSP webinar this month entitled "Analyzing Customer Data to Create Competitive Advantage".   I have been a 'numbers guy' since early in my career at A.C. Nielsen, so I understand how powerful data can be if transformed into actionable information. There are mountains of data which can be analyzed, but In order to find unique insights, you need to not only access to the information, but also the tools which allow you to manipulate, store, and create custom data sets in real time.  

Analytics touches many aspects of scholarly publishing; content creation, editorial, marketing and sales, product packaging/pricing, and audience development.  Typical business questions can include most read articles, turnaway reports, usage for a specific event such as an annual conference, or simply tracking your content's overall performance. For my section of the webinar, I plan on examining a few of these potential questions and how to extract "valuable nuggets" from data analysis which can provide powerful editorial and marketing insights.

I hope many of you can make it on November 15th!

Bill










For more details on the webinar and to sign up, visit the ALPSP website.  If you can't make the 15th, make sure you sign up to get access to the webinar later.


Bill Matthews
Director of Business Development
HighWire Press | Stanford University
(o) 650-725-9279
(m) 650-208-4643
bmatthews@highwire.stanford.edu

HighWire:  Pathfinders in Scholarly Publishing