Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Friday, 19 May 2017

Not Your Teenager’s Social Network: What Academic Societies Can Learn from Facebook about Making Money—and Making Members Happy



We are delighted to share this blog by Roy Kaufman, Managing Director of New Ventures at the Copyright Clearance Center  which draws some interesting parallels between learned societies and social networks and highlights what we can learn from their example.


image social media networkingSocial networking a la Facebook and Twitter may seem to be a product of the Internet age, but it is actually nothing new. If you think about it, learned societies - which aim to bring together people in a given field or area of professional interest - are actually built on the original idea of a social network, to wit: a network of social interactions and personal relationships, as Webster’s defines it.

Yet today’s academic societies, charged with connecting individuals who share professional interests and providing a forum for communication and collaboration, continuing education, and career opportunities, are facing declines in membership, particularly among people under age 30. Fewer than half (48%) of all millennials belong to a society compared with 83% of baby boomer researchers, according to Wiley’s recent survey of nearly 14,000 research professionals.

Facebook and Twitter (and more researcher-focused sites such as Mendeley) have something to do with that age discrepancy; they are favorites for early-career researchers who want to actively network in both their professional and personal life. With so many online opportunities for making contacts and interacting, societies are tasked with finding new ways to provide meaningful benefits that will attract and retain members, as well as keep their revenues growing.

Perhaps scholarly and professional societies can learn something from Facebook, too. Just as that online social network continues to expand (and gobble up money) by using member data in ever more ingenious ways (linking all those Likes, learning from them, and tailoring content to members accordingly), so societies can use technology to better serve their members and become more relevant and profitable in the process.

From disconnected data to smart data

Fully leveraging data they already have is an often-overlooked way for societies to grow their membership and keep current members engaged enough to renew year after year. Take the example of researchers who submit an article to a society journal. It is a good bet that the article will contain the names and contact information of multiple coauthors. Wouldn’t it make sense if, instead of isolating those names within the editorial system, societies could use them to their advantage, connecting them with other data points throughout the organization?

That process could start at article submission by determining the needs of corresponding authors and co-authors, simply by asking questions like the following:

  • Are the authors already members of the society?
  • If yes, are their memberships up for renewal?
  • If no, will a special offering (such as an APC discount or free author reprints) entice them to join?
  • If they have an .edu address, are they taking advantage of institutional arrangements for payment of open access fees?
  • Are they registered to attend the next society conference?
  • Do they need continuing education credits?
  • Do they even know about these benefits?

Chances are, members and would-be members don’t know all the benefits of membership. In the Wiley survey, 15% of respondents said they’d never been invited to join an academic society; 12% said they didn’t know what offerings were available, and another 12% said that joining had never occurred to them. As the folks at Wiley put it, “This means that 37% of non-members are either waiting to be asked to join, or might be persuaded to join…With so many non-members just waiting to be asked, societies may find they are often pushing at an open door.”

But societies are not yet pushing on that door. One reason is that in many learned societies, different departments, and the data they house, are “siloed,” cut off from one another and not communicating effectively. “When a member interacts with an organization, they’re interacting with education, or a group that does grants,” says Ann Michael, DeltaThink CEO and former president of the Society for Scholarly Publishing, who moderated a recent webinar on society membership by the Copyright Clearance Center. Silos, says Michael, make it difficult for members to see the organization as a whole, which makes it hard for organizations to serve members’ needs effectively.

In the same webinar, Alex Taylor, head of communities and events at the Institution for Engineering and Technology, admitted that for a long time, the IET was “lost in a labyrinth of our own making .…We offer so many different things, [there are] so many different teams and departments…that there’s most definitely a [silo] culture, a lack of joined-up collaborative thinking.”

The key, then, is for members and societies to come together, to increase satisfaction and engagement on one side and revenues on the other. That starts with knowing what members and potential members want and need.  For example, in Wiley’s survey findings, 26% of respondents said their strongest reason for joining a society was to take advantage of opportunities for continuing education. But the continuing education platforms seldom, if ever, talk with the editorial ones.

The bottom line: If the membership, continuing education, and conference departments are not connected with each other or linked up with the editorial department, opportunities for generating new members and retaining existing ones will be missed. Think about the benefits to all involved if these systems talked to one another. In that scenario, it would be easy to notify an individual who recently submitted an article on a particular topic about an upcoming workshop on the same subject. Or, having just published that article, to let the author know that his society membership renewal comes with the benefit of 25 free article reprints.

What all of this requires is a smart network of links among databases that enables societies to target their marketing to specific individuals with personalized messages and offerings, at opportune times (when you already have their attention, for example, at article acceptance or other key points in the editorial workflow). That is the difference between blasting members with renewal notices three days after they’ve renewed and instead telling them something they truly want to know (i.e. that they are due for CME credits). Rather than putting off members and would-be members with more junk mail, suddenly, you are providing them with a higher level of service.

One way to make the data connection easy is with an enterprise content management system that does the sorting and linking of member information automatically. An investment in an ECM system is worth it, because it allows societies to provide a higher level of service.  Knowing what members need and offering it to them when they need it will bring in higher revenues in the form of new and renewing members, who can now avail themselves of services they were previously unaware of. Or, to put it another way, societies will be able to maximize revenue sources already at their disposal, and members will understand the value proposition that comes from joining and engaging with a learned society. Talk about pushing an open door.

Adopt some standards

Besides enterprise content management systems, another crucial step toward connecting data and better serving members is to adopt standards such as ORCID IDs, Ringgold names, IP addresses from Publisher Solutions International, and identifiers from FundRef. Once employed, societies can identify institutional affiliations, funding agencies, geographical locations, and membership status, and then launch relevant messaging.  You might, by ORCID ID, identify an author member as hailing from a particular institution and take it from there, reaching out to let a Harvard-based author know that she’s eligible for an institutional discount on open access charges. Combine these standards with the member data derived from your enterprise content management system, and suddenly, you get to the nirvana of data connection, without having to reinvent the wheel, and without having to bother the author.

Create new businesses to keep members happy

To keep growing, societies also need to consider new sources of revenue. It makes sense that the first thing a membership-driven society should consider when it thinks about growing its bottom line is the needs of its members. For example, the Wiley survey asks members what they value.  Some key services mentioned in the Wiley survey are continuing education (64%), keeping up to date with the latest research (50%) and job openings (32%). Once societies have this information in hand, they should ask: Do I have a business around this? If the answer is no, the next question might be: Should I have a business around this? If learning is a key reason members renew, a society may want to look at whether they have adequate continuing education offerings. If career networking is a top priority, a society might send out alerts when jobs open up in members’ areas of interest. That’s known as data driven messaging, whether a society tells a researcher who has just submitted an article on kidney cancer about an opening in the nephrology department of a major research hospital, or reminds her to register for the upcoming American Society of Nephrology Conference.

Attracting and retaining members - even millennials - is not rocket science, and we can learn from the companies who do it well. It is about figuring out why people join, and asking: Have I done enough here? Because sometimes, by asking relatively simple questions, offering opportunities vis-à-vis the needs of members, and doing some obvious things like adopting standards, it is possible to create the building blocks that raise a society to the next level - and make it go viral. 

photo Roy Kaufman
Roy Kaufman is Copyright Clearance Center's Managing Director of New Ventures. Prior to CCC, Roy served as Legal Director, Wiley-Blackwell, John Wiley and Sons, Inc. He is a member of, among other things, the Bar of the State of New York, the Copyright and Legal Affairs Committee of the International Association of Scientific Technical and Medical Publishers. He was the founding corporate Secretary of Crossref, and formerly chaired its legal working group. He has lectured extensively on the subjects of copyright, licensing, open access, text/data mining, new media, artists’ rights, and art law. Roy is Editor-in-Chief of Art Law Handbook: From Antiquities to the Internet, and author of two books on publishing contract law. He is a graduate of Brandeis University and Columbia Law School.

Make sure you get the most out of your ALPSP membership? Visit our Membership Benefits page to keep up to date on all our services on offer. For any queries please contact Lesley Ogg at events@alpsp.org 

Wednesday, 14 September 2016

Plenary 1: The Conversation: Research and Scholarly Publishing in the Age of Big Data

Ziyad Marar is Global Publishing Director at SAGE Publishing. Chairing the first plenary session of the ALPSP conference, he engaged his colleague Ian Mulvany, Head of Product Innovation, and Fran Bennett, CEO and co-founder of a big data company Mastodon C in a conversation about publishing in the age of big data.

Is big data hype and nonsense - just an exciting term that let's an agency sell their services? Fran Bennett believes there are some fundamental things that have changed that mean it is so much more than that. It can help companies open up new insights, generate additional income and lower barriers to technology entry. As the technology gets better it can do different applications. There is more data and cheaper processing.


Mastodon C are working with the UK Government department responsible for animals and farming. They are collecting all the data of dead livestock. They don't have enough staff so sometimes patterns get missed. They use computers to identify any of these threads to analyse post mortem. They can take messy structural data and sorts it out so expert humans can use their time more effectively and in a targeted way.

Ian Mulvany thinks high quality content is what we do as an industry, but it's all digitally mediated content. All publishing organizations need to be technologically competent. We're in a mixed world of software solutions that are beginning to be commodified. But the variety of the services around them are living in a handwritten world: a dilemma he is endlessly fascinated by.

Corporate applications of big data can transfer to publishing in market projections, customer retention, internal SWOT analysis and with hiring. Mulvany asks how many publishers have tried to re-analyse their entire corpus using big data techniques? Not many hands went up... there are lots of opportunities here. Bennett observed that a good data scientist is a statistician who can code and understand the context of their data and warned against tracking things purely because you can: the risk is you create 'data exhaust' that you can't do anything with.

Mulvany noted that some fields have long worked with big data and have good standards and procedures to deal with it. He is particularly interested in working with researchers that have realised they have a whole load of data and don't know what to do with it. There is a 'data under the desk' problem. Data is collected sporadically, is not necessarily kept well, and isn't large scale.

Caution was called for by delegates in the audience and on Twitter when using algorithms for peer review: it can and will be exploited by researchers. The panellists all agreed that machines can do the dirty work for us, but not all the work.

Marar outlined the work of the Berkeley sociologist, Nick Adams, who is using crowdsourcing and algorithms to look at reports on the Occupy movements in nine cities. Analysis that would normally have taken 15 years has actually taken one year, and is finding interesting patterns. He also cited the work of Gary King, a Harvard social scientist who is developing and applying empirical methods in many areas of social science research, focusing on innovations that span statistical theory to practical application.

Social researchers are coming more slowly to big data analysis, but are doing some unusual work with it. SAGE Publishing has conducted a massive survey into the area of data and social science with over 13,000 responses. It's something they are focusing on as a priority.

An interesting side issues when looking at social data is sometimes, when you look at the data, you find that the quality of it is not what it might be, with potential to lead to data protection breaches on a grand scale. There are differences between ethical and legal behaviour concerning datasets. it may be cheap to capture and hold data, but expensive to extract, clean and deliver it.

Mulvany closed with the observation that there are researcher needs, potential development tools, but why should the industry care about these things? Because at our heart we are about democratising knowledge and finding the right solutions and people around that knowledge. If we look purely at their purpose it will give us the realisation on how we make it happen. Those tools are becoming cheaper to experiment and innovate with. So we should do so.

Ziyad Marar is Global Publishing Director at SAGE Publishing where Ian Mulvany is Head of Product Innovation. Fran Bennett is CEO and Co-Founder of Mastodon C. They took part in a panel discussion at the ALPSP Conference 2016.

Monday, 18 May 2015

High Value Content: Big Data Meets Mega Text

ALPSP recently updated the Text and Data Mining Member Briefing (member login required). As part of the update, Roy Kaufman, Managing Director of New Ventures at Copyright Clearance Center, provided an overview of the potential of TDM, outlined below.

"Big data may be making headlines, but numbers don’t always tell the whole story. Experts estimate that at least 80 percent of all data in any organization—not to mention in the World Wide Web at large— is what’s known as unstructured data. Examples include email, blogs, journals, Power Point presentations, and social media, all of which are primarily made up of text. It’s no surprise, then, that data mining, the computerized process of identifying relationships in huge sets of numbers to uncover new information, is rapidly morphing into text and data mining (TDM), which is creating novel uses for old- fashioned content and bringing new value to it. Why? Text-based resources like news feeds or scientific journals provide crucial information that can guide predictions about whether the stock market will rise or fall, can gauge consumers’ feelings about a particular product or company, or can uncover connections between various protein interactions that lead to the development of a new drug.

For example, a 2010 study at Indiana University in Bloomington found a correlation between the overall mood of the 500 million tweets released on a given day and the trending of the Dow Jones Industrial Average. Specifically, measurements of the collective public mood derived from millions of tweets predicted the rise and fall of the Dow Jones Industrial Average up to a week in advance with an accuracy approaching 90 percent, according to study author Johan Bollen, Ph.D., an associate professor in the School of Informatics and Computing. At the time, Dr. Bollen predicted, with uncanny accuracy, where he felt TDM was going, from the imprecise, quirky world of Facebook and Twitter to high-value content. He said, "We are hopeful to find equal or better improvements for more sophisticated market models that may in fact include other information derived from news sources and a variety of relevant economic indicators."

In other words, structured data alone is not enough, nor is text mined from the wilds of social media. Wall Street and marketers, eager to predict the right moment to hit buy or sell or to launch an ad campaign, have already moved from mining Facebook and Twitter to licensing high-value content, such as raw newsfeeds from Thomson Reuters and the Associated Press, as well as scientific journal articles reformatted in machine- readable XML. In fact, a 2014 study by Seth Grimes of Alta Plana concludes that the text mining market already exceeds 2 billion dollars per year, with a CAGR of at least 25%.

Far from being irrelevant in our digital age, high-value content is about to have its moment, and not just to improve the odds in the financial world or help marketers sell soap. It represents a new revenue stream for publishers and their thousands of scientific journals as well. For example, in 2003, immunologist Marc Weeber and his associates used text mining tools to search for scientific papers on thalidomide and then targeted those papers that contained concepts related to immunology. They ultimately discovered three possible new uses for the banned drug. “Type in thalidomide and you get between 2,000 and 3,000 hits. Type in disease and you get 40,000 hits,” writes Weeber in his report in the Journal of the American Medical Informatics Association. “With automated text mining tools, we only had to read 100-200 abstracts and 20 or 30 full papers to create viable hypotheses that others could follow up on, saving countless steps and years of research.”

The potential of computer-generated, text-driven insight is only increasing. In his 2014 TedX Talk, Charles Stryker, CEO of the Venture Development Center, points out that the average oncologist, after scouring journals the usual way, reading them one by one, might be able to keep track of six or eight similar cancer cases at a time, recalling details that might help him or her go back, re-read one of two of those articles, and determine the best course of care for a patient with an intractable cancer. The data banks of the two major cancer institutes, on the other hand, hold searchable records of cancer cases that can be reviewed in conjunction with 3 billion DNA base pairs and 20,000 genes contained within each. So using that data would mean a vast improvement in the odds of finding clues to help treat a tricky case or target the best clinical trial for someone with a rare disease. This information might otherwise have been difficult, if not impossible, for even the most plugged-in oncologist to find, let alone read, see patterns, or retain the information for a period of time.

Think, then, of the possibilities of improving healthcare outcomes if the best biomedical research were aggregated in just a few, easily accessible repositories. That’s about to happen. My employer, Copyright Clearance Center (CCC), is coming to market with a new service designed to make it easier to mine high-value journal content. Scientific, technical and medical publishers are opting into the program, and CCC will aggregate and license content to users in XML for text mining. Although the service has not yet fully launched, CCC already has publishers representing thousands of journals and millions of articles participating.

Consider the difficulties of researchers, doctors, or pharmaceutical companies wishing to use text mining to see if cancer patients on a certain diabetes drug might have a better outcome than patients not on the drug. They must go to each publisher, negotiate a price for the rights, get a feed of the journals, and convert that feed into a single useable format. If the top 20 companies did this with the top 20 publishers, it would take 400 agreements, 400 feeds, and 400 XML conversions. The effort would be overwhelming.

Instead, envision a world where users can avail themselves of an aggregate of all relevant journals in their field of interest. Instead of 400 agreements and feeds to navigate and instead of 400 documents to convert to XML, there would be maybe 40 agreements: 20 between the publishers and CCC and 20 with users. There would be no need for customers to convert the text. In other words, researchers could get their hands on the high-value information they need to move research and healthcare forward, in less time, with less effort. And that’s only the beginning. As Stryker said about the promise of TDM, “We are in the first inning of a nine-inning game. It’s all coming together at this moment in time.”

ALPSP Members can login to the website to view the Briefing here.

Roy Kaufman is Managing Director of New Ventures at the Copyright Clearance Center. He is responsible for expanding service capabilities as CCC moves into new markets and services. Prior to CCC, Kaufman served as Legal Director, Wiley-Blackwell, John Wiley and Sons, Inc. He is a member of the Bar of the State of New York and a member of, among other things, the Copyright Committee of the International Association of Scientific Technical and Medical Publishers and the UK's Gold Open Access Infrastructure Program. He formerly chaired the legal working group of CrossRef, which he helped to form, and also worked on the launch of ORCID. He has lectured extensively on the subjects of copyright, licensing, new media, artists' rights, and art law. Roy is Editor-in-Chief of ‘Art Law Handbook: From Antiquities to the Internet’ and author of two books on publishing contract law. He is a graduate of Brandeis University and Columbia Law School.


Friday, 17 October 2014

Mind the (data) gap… Learned Publishing special issue

Fiona Murphy (centre) talks Data at the ALPSP conference
For anyone who has ever travelled on the London Underground and endured endless repeats of ‘Mind the gap’, this special issue of Learned Publishing is your equivalent warning on data.

As funders make open data a policy stipulation, publishers must prepare for these requirements. In fact publishers are well placed to support open data, and society publishers are uniquely well placed to be a part of the solution: they are at the heart of their community and understand their needs.

But what do you do next? How can you mind your data gap and understand what it means for your organization and its community?

In this special online-only issue of Learned Publishing, the focus is purely on data. Guest edited by Alice Meadows, Director of Communications at Wiley and Fiona Murphy, STM Publisher, it is published open access with the support of Wiley.

We caught up with Alice and Fiona (who was just back from last month's European Research Council Workshop on Research Data Management and Sharing in Brussels), to talk data deluge and why now for this special issue.

So why focus on data now?


Alice: The OSTP memo from 2013 and the European Commission’s Horizon 2020 Research Data Pilot are two examples of funders driving open data. Meanwhile, more data than ever are being collected. Technology is improving our ability to analyse and share them, but there are still huge barriers to that being done effectively; you’d be surprised how much data collection is still manual.

Fiona: And we still lack globally standard ways of collecting, managing, sharing, and storing data which creates a whole new set of challenges when the ultimate aim is to enable re-use and interoperability. Susan Reilly's paper provides a librarian's perspective of some of these issues, while Varsha Khodiya and her F1000 colleagues tackle data sharing, citation, and more.

What can publishers to do help?


Alice: If publishers and societies aren't careful, they will once again be playing catch-up with the funders on a growing requirement in their research communities. This is a golden opportunity to lead from the front and help researchers. In the words of Mark Hahnel in this interview on the Wiley Exchanges blog, open data can help “Opening up research data has the potential to both save lives (say with medical advances) and to enhance them with socio-economic progress.” That’s a pretty compelling argument. And societies and society publishers have a particular part to play here, as demonstrated in the paper by Hazel Norman of the British Ecological Society. 

Fiona: That’s not to say that publishers aren't already working on opening up data. My paper Data and Scholarly Publishing: the transforming landscape sets the scene and provides an overview of how publishers are responding to date.

ALPSP: What is the most important theme to emerge from the issue?


Fiona: Without a doubt, it’s the importance of collaboration. Cooperation between stakeholders is crucial to successfully opening up data. Andrew Treloar reflects on the work of the Research Data Alliance in his paper. Having recently returned from a European Research Council workshop attended by a whole cross-section of stakeholders, I can only agree that these types of coordinated action are the best way forward.

Alice: Similarly, Sarah Callaghan's paper on preserving the integrity of the scientific record shows how the scholarly community is collaborating to solve issues around data citation and linking. But if they’re not familiar with recent developments or with networks like RDA, I’d urge readers to access the articles, share with colleagues and talk through what it means for their organization. This special issue is a snapshot of views from right now - things are likely to change rapidly. We’d love to know what ALPSP members think and if they have positive examples and experiences they can share.

Learned Publishing special data issue is available now online open access on the Learned Publishing site.


Friday, 26 September 2014

The data deluge is upon us… are you ready?

Today sees the online publication (under an open access model) of our special data issue of Learned Publishing.

Produced with the support of Wiley, this collection of papers represents a snapshot of current thinking about research data from a variety of perspectives. It is guest edited by Alice Meadows, Director of Communications at Wiley and Fiona Murphy, their STM Publisher.

This special issue is launched at the end of a week where Wiley Exchanges published some fascinating posts on different aspects of data to coincide with the Research Data Alliance annual conference in Amsterdam.

Liz Ferguson, Publishing Solutions Director at Wiley hit the nail on the head with her observation “Acknowledging the significance of data in scholarly communication is one thing, but knowing what to do about it is another” in her piece Everybody Loves Data.

Jennifer Beal, Events & Ambassador Manager at Wiley observed “Ah Big Data, how things have changed!” in her write up of the Who’s Afraid of Big Data session from the ALPSP conference.

“Do you want to use my environmental data, or I yours? The question pulls in many conflicting directions.” Mike Kirkby, Emeritus at Leeds University reflected on the many questions with complex answers that the use and storage of data presents in  More data, more questions?

Fiona Murphy interviewed Mark Hahnel, Founder of figshare who believes that “Opening up research data has the potential to both save lives (say with medical advances) and to enhance them with socio-economic progress. It’s a space where humans and computers can work symbiotically, and where industry can also benefit.” He goes on to share his thoughts on the practicalities of opening data, blockages in the system and the potential for open science.

As more and more colleagues across the scholarly publishing community engage with open data, we hope this special issue will help them along the way.