Showing posts with label Fiona Murphy. Show all posts
Showing posts with label Fiona Murphy. Show all posts

Thursday, 4 February 2016

All Change in Scholarly Communications: How are the Players – Veterans and Newbies – Adapting?

Fiona Murphy reports from #APE2016
Last month, in characteristically bracing January Berlin weather, around 250 intrepid speakers and delegates attended the 11th Academic Publishing in Europe (APE – pronounced “Ahhhpay”) meeting. Keep an eye on Twitter #ape2016 as all of the presentations were recorded and so should become available in the near future.

A number of familiar characters – large publishers, established platform providers, and so forth – whose language seems to have evolved over the past few years – spoke about ‘openness’ and ‘sharing’ rather than preserving business models. Todd Toler of Wiley, for instance, expressed the “publisher’s value proposition” as having shifted from content provision – basically “moving stuff about” to “strengthening knowledge connections”. This feels like a real turning of tides; such players are now actively aiding and abetting our efforts to garner significant knowledge from our scholarly ecosystem.

In point of fact, there was a general theme around intelligence rather than simply the power of data. Barend Mons bemoaned the existence of “a Christmas tree of hyperlinks and the malpractice of supplementary material’”, instead calling for the training of experts to really understand how machine learning and human interrogation of data can be meshed together to form a powerful whole – “Open Science as a Social Machine” (keep an eye on the IDCC programme in Amsterdam later this month, as he’ll be expanding on the topic there). Meanwhile, Emma Green, of Zapnito – a start-up that aids knowledge-based companies to maximise the impact of their associated experts spoke of growing the ‘knowledge economy’ by reducing the noise and chatter, thereby freeing up the collective intelligence.

John Sack of Highwire’s approach was to examine frictions in the workflow. If workflow is ‘a way of getting things done,’ then instances of friction – with the possible exception of a review stage – largely involve the loss of efficiency. Currently most journal workflows are still based on the original print journal format, but with the version of record shifting online, the resulting misalignments between what is desired and what is produced are causing delays, and infringements of established rules (such as copyright). Friction-reducing tools that can support and simplify the generation, finding, and attribution of scholarly outputs are needed. This can be enabled by standards such as e.g. ORCID or ResearcherID for people, and by initiatives such as openRIF/VIVO for connecting people and their roles to their works and activities. This connectivity will surely boost quality, productivity, and the need for improved garnering of knowledge from our research landscape that generally arose as a theme across APE in general. This connectedness, according to Sack, is about a supported conversation amongst collaborators who are enabled by tools that sift, pre-curate and – potentially – publish their scholarly outputs.

Opportunities for new business models are appearing in a number of points in the workflow – Publons acknowledges and badges peer review activities, Overleaf provides templated support to write journal articles, and Elsevier is leveraging the new Mendeley Data service to enable authors to publish their data and link it immediately with journal articles.

At the same time, policy (=funding) is also moving in the same direction. Stephan Kuster, Head of Policy Affairs for Science Europe explained its function and mission. Science Europe is a think tank set up to support and advise EU National Research Funding Councils around on EU R&D policy issues. Open Access is one of nine key priorities, including enabling authors to hold copyright, supporting sustainable archiving, and publication and dissemination are integral part of research process and should be funded as such.

There was a thoughtful debate about Scholarly Communications Networks and whether they add value, which would not have been possible even a few years ago. Fred Dylla, Emeritus Executive Director of the American Institute of Physics, made the salient point that reputation of the journal still needs to be fundamentally challenged for the landscape to be really disrupted. Currently, the people and institutions making the key decisions about funding, tenure and promotion, are still fixated on journal reputations and impact factors. So, despite feeling as though there has been a lot of progress in the last few years, it also seems there’s still a lot to do.

Luckily there are several opportunities coming up to extend and develop our understanding of and strategies for adapting to this changing landscape. As well as the aforementioned IDCC later this month. And look out for the ALPSP Seminar on research data, digital preservation and innovation in March. Standing on the Digits of Giants is co-organised with the Digital Preservation Coalition and is designed to orientate and empower publishers, research managers and researchers to navigate and flourish in the new landscape.

Another key space to continue these discussions is in the context of the Force11 community, which aims to bring together many of the stakeholders needed at the table to effect change: policy makers, funders, researchers, technologists, publishers, informaticists, lawyers, etc. Force16 promises to be an exciting venue where we’ll be pushing scholarly communications into uncharted territory. Hope to see you there too.

Fiona Murphy, February 2016

Now associated with the Maverick Publishing Specialists, Fiona Murphy has held a range of production and editorial roles at Wiley, Oxford University Press, Random House and Bloomsbury Academic. She specializes in emerging scholarly communications (including Open Science and Open Data) and works to raise expertise and activity levels across the wider research and publications communities. Fiona has written and presented extensively on the research landscape, data and publishing. She is Co-Chair of the World Data System—Research Data Alliance Publishing Data Workflows Working Group, an Editorial Board Member of the Data Science Journal and enjoys organizing meetings. orcid.org/0000-0003-1693-1240

This post was written by Fiona Murphy with the support of Melissa Haendel.



Friday, 17 October 2014

Mind the (data) gap… Learned Publishing special issue

Fiona Murphy (centre) talks Data at the ALPSP conference
For anyone who has ever travelled on the London Underground and endured endless repeats of ‘Mind the gap’, this special issue of Learned Publishing is your equivalent warning on data.

As funders make open data a policy stipulation, publishers must prepare for these requirements. In fact publishers are well placed to support open data, and society publishers are uniquely well placed to be a part of the solution: they are at the heart of their community and understand their needs.

But what do you do next? How can you mind your data gap and understand what it means for your organization and its community?

In this special online-only issue of Learned Publishing, the focus is purely on data. Guest edited by Alice Meadows, Director of Communications at Wiley and Fiona Murphy, STM Publisher, it is published open access with the support of Wiley.

We caught up with Alice and Fiona (who was just back from last month's European Research Council Workshop on Research Data Management and Sharing in Brussels), to talk data deluge and why now for this special issue.

So why focus on data now?


Alice: The OSTP memo from 2013 and the European Commission’s Horizon 2020 Research Data Pilot are two examples of funders driving open data. Meanwhile, more data than ever are being collected. Technology is improving our ability to analyse and share them, but there are still huge barriers to that being done effectively; you’d be surprised how much data collection is still manual.

Fiona: And we still lack globally standard ways of collecting, managing, sharing, and storing data which creates a whole new set of challenges when the ultimate aim is to enable re-use and interoperability. Susan Reilly's paper provides a librarian's perspective of some of these issues, while Varsha Khodiya and her F1000 colleagues tackle data sharing, citation, and more.

What can publishers to do help?


Alice: If publishers and societies aren't careful, they will once again be playing catch-up with the funders on a growing requirement in their research communities. This is a golden opportunity to lead from the front and help researchers. In the words of Mark Hahnel in this interview on the Wiley Exchanges blog, open data can help “Opening up research data has the potential to both save lives (say with medical advances) and to enhance them with socio-economic progress.” That’s a pretty compelling argument. And societies and society publishers have a particular part to play here, as demonstrated in the paper by Hazel Norman of the British Ecological Society. 

Fiona: That’s not to say that publishers aren't already working on opening up data. My paper Data and Scholarly Publishing: the transforming landscape sets the scene and provides an overview of how publishers are responding to date.

ALPSP: What is the most important theme to emerge from the issue?


Fiona: Without a doubt, it’s the importance of collaboration. Cooperation between stakeholders is crucial to successfully opening up data. Andrew Treloar reflects on the work of the Research Data Alliance in his paper. Having recently returned from a European Research Council workshop attended by a whole cross-section of stakeholders, I can only agree that these types of coordinated action are the best way forward.

Alice: Similarly, Sarah Callaghan's paper on preserving the integrity of the scientific record shows how the scholarly community is collaborating to solve issues around data citation and linking. But if they’re not familiar with recent developments or with networks like RDA, I’d urge readers to access the articles, share with colleagues and talk through what it means for their organization. This special issue is a snapshot of views from right now - things are likely to change rapidly. We’d love to know what ALPSP members think and if they have positive examples and experiences they can share.

Learned Publishing special data issue is available now online open access on the Learned Publishing site.


Friday, 26 September 2014

The data deluge is upon us… are you ready?

Today sees the online publication (under an open access model) of our special data issue of Learned Publishing.

Produced with the support of Wiley, this collection of papers represents a snapshot of current thinking about research data from a variety of perspectives. It is guest edited by Alice Meadows, Director of Communications at Wiley and Fiona Murphy, their STM Publisher.

This special issue is launched at the end of a week where Wiley Exchanges published some fascinating posts on different aspects of data to coincide with the Research Data Alliance annual conference in Amsterdam.

Liz Ferguson, Publishing Solutions Director at Wiley hit the nail on the head with her observation “Acknowledging the significance of data in scholarly communication is one thing, but knowing what to do about it is another” in her piece Everybody Loves Data.

Jennifer Beal, Events & Ambassador Manager at Wiley observed “Ah Big Data, how things have changed!” in her write up of the Who’s Afraid of Big Data session from the ALPSP conference.

“Do you want to use my environmental data, or I yours? The question pulls in many conflicting directions.” Mike Kirkby, Emeritus at Leeds University reflected on the many questions with complex answers that the use and storage of data presents in  More data, more questions?

Fiona Murphy interviewed Mark Hahnel, Founder of figshare who believes that “Opening up research data has the potential to both save lives (say with medical advances) and to enhance them with socio-economic progress. It’s a space where humans and computers can work symbiotically, and where industry can also benefit.” He goes on to share his thoughts on the practicalities of opening data, blockages in the system and the potential for open science.

As more and more colleagues across the scholarly publishing community engage with open data, we hope this special issue will help them along the way.

Thursday, 11 September 2014

Who's afraid of big data?

Who's afraid of big data? panel
Fiona Murphy from Wiley chaired the final panel on day two of the 2014 ALPSP International Conference. She posed the question: how do we skill up on data, take advantage of opportunities and avoid the pitfalls?

Eric T. Meyer, Senior Research Fellow and Associate Professor at the University of Oxford was first up trying to answer. He observed how a few years ago you would struggle to gain an audience for a big data seminar. Today, it's usually standing room only.

Big data has been around for years. People were quite surprised when Edward Snowden leaked the NSA documents via Wikileaks, but it had been going on for a long time. Big data in scholarly research has also been around a long time in certain disciplines such as physics or astronomy. There was always money to be made in big data, but there's even more now, and everyone is starting to realise it. So much so, you need a big data strategy.

Meyer defines big data as data unprecedented in scale and scope in relation to a given phenomenon. It is about looking at the whole datastore rather than one dataset. Big data for understanding society is often transactional. We're talking really big. If you can use it on your laptop, it won't be big data.

Meyer drew on some entertaining examples of how big data can be used. If you key in the same sentence in different country versions of Google you'll see the variety of responses change. There are limits to big data approaches, they can come up with misleading results. What happens when bots are involved? Does it skew the results? The challenge will be how you can make it meaningful and more useful.

David Kavanagh from Scrazzl reflected on how the challenge researchers face when making decisions about how to structure and plan your experiments. If you want to leverage collective scientific knowledge and identify which products you want to use for your work, there wasn't a structured way of searching of doing this. Kavanagh urged publishers to throw computational power at data and content as a way to solve problems, improve how you work and help make sense of unstructured content.

That's what they have tackled with Scrazzl which is a practical application of structured or unstructured data that Eric Meyer mentioned. You need to have a product database. Then you have to cut out as much human intervention out as you can. Automation is key. Where they couldn't find a unique identifier or a catalogue match, they had to make it as fast as possible for a human to make an association. Speed is key.

Finally, they built a centralised cloud system that vendors could update their own records. It's a crowd sourced system for those who have a vested interest in keeping it up-to-date. The opportunity for them going forward will be through releasing this structured information through unstructured APIs to drive new insights. It also allows semantic enablement of content and offers the opportunity to think about categorisation in new ways.

For publishers running an ad supported model, they can get use the collection of products from the content search and then identify which advert is the most suited for you.

Paul F. Uhlir
Paul F. Uhlir  from the Board on Research & Information at The National Academies  observed that even after 20 years, we have yet to deconstruct the print paradigm and reconstruct it on the Net very well. In the 1980s a gigabyte was considered a lot of data. in the 1990s. a terabyte was a lot of data. In this decade, the exabyte era is not far ahead of us and a whole lot of others ahead of it.

Huge databases in business, mining marketing information and other data. The Internet of Things and semantic web. Everything now can be captured, represented and manipulated in a database. It's an issue of quantity. But there is also an issue of quality. There needs to be a social response. There are a series of threats around big data.



Disintermediation

The rise of big data promises a lot more disruption. Think about 3D printing. The consequence could be millions of product designers specifying items. Manufacturing will be affected. Jobs will be lost. What will happen to the workers in a repair and body shop when cars are driverless? What will happen to the insurance industries. Workers will be disintermediated. What is certain is that there will be massive labour shifts and disruptions.

Playing God

Custom organs for body parts. The ability of insert genes into another organism. All these applications are data intensive and will become even more so. They have profound social and ethical issues and have potential to do great harm.

End of Privacy

Meyer touched on the NSA files. What about spying satellites? The ubiquity of CC TV? These images are kept in huge databases for future use. Product information is held and used to identify preferences by private companies. There is no such thing as privacy any more.

Inequality

Big data are increasingly powerful means to increase hold on global power.

Complexity

The more we learn, the less we know. Any scientist will tell you that greater understanding leads to more questions than answers.

Luddite reactions

The reaction of people to the encroachment of strange and frightening techniques of the technology age where through passive resistance they try to lead a simple life.

There are also a number of weaknesses that centre around:

  • Improving the policies of the research community
  • New or better incentive mechanisms versus mandates
  • Explicit links of big data to innovation and tech transfer
  • Changing legal landscape-lag in law/bad law/IP law
  • Data for policy-communicating with decision makers.




Sunday, 15 September 2013

Data: not the why, but the how (and then, what?)

Simon Hodson introduces CODATA
It is now old news that data - its production, management and re-use potential - is of growing significance to all the key stakeholders within the scholarly communication ecosystem. Publishers need to navigate the emerging landscape of technology, researcher and industry needs, and funders’ and policy-makers’ priorities in order to continue supporting the growth and discoverability of knowledge.

Wiley's Fiona Murphy chaired a panel discussion that provided publishers with insights on how – and by whom – the roadmap is being written, as well as its challenges and potential opportunities from researchers, funders and industry perspectives.

Simon Hodson, Executive Director at CODATA, provided an overview of their work. Their focus is on strengthening international science for the benefits of society by promoting improved scientific and technical data management and use. It is an international community and network of expertise on data issues. Their key areas of activity are policy frameworks for data, frontiers in data science and technology and data strategies. They develop data citation, standards and practices. September sees the release of their major report Out of Cite, Out of Mind.

Hodson provided an overview of relevant data policy including the Royal Society Science as an Open Enterprise Report from 2012. Examples of projects relating to data policy include the Dryad Joint Data Archiving Policy and Dryad Sustainability.

Kerstin Lehnert, Director of the Integrated Earth Data Applications Research Group (IEDA), provided a view of data from the researcher's perspective. Why open access to data? There are two main reasons: to allow verification of research results and to make data accessible for re-use.

Kerstin Leynert on useful data
Data must be fit for re-use: it must be discoverable, be openly accessible, be safe and it must be useful. One example of useful data is the Sloan Digital Sky Survey. It contains 2000 articles with over 70,000 citations, and a lot of useful science has come out of it. Another example is Earthchem Synthesis. Within 2 minutes you can explore the whole literature and create a map with different composition of different areas. It has seamless integration within the discipline.

There are a number of guiding principles. Quality of data makes it useful to include complete documentation of provenance. Domain-specific data stewardship provides development, maintenance and promotion of domain-specific, community based standards for data and metadata. Domain-specific repositories are best positioned to ensure 'fitness for re-use'. However, they must ensure professional data curation services and integrate with the scholarly communication ecosystem.

Lehnert outlined the IEDA development of standards. They had requirements for the reporting of geochemical data. Steps were taken to move from a suite of databases to become a repository to ensure better sustainability. They improved policies and procedures using DOI, IGSN and long term archiving agreements with NGDC and Columbia University libraries. They also sought accreditation with membership in the world data system and as a publication agent of DataCite.

Lehnert closes with a number of questions. Many data types have no home or standards. Where should these go? How can we help other domains to establish repositories and best practices? Who decides which repository data should be submitted to? Are there recommendations from societies? The sustainability of repositories is still not solved, so what are the business models to ensure longevity? How can we streamline the link between repositories and journals? Do we need a centralised solution?

The final speaker was Tony Brookes from the Department of Genetics at the University of Leicester. He urged data sharing as it is important, but acknowledged it is problematic. Prioritise IDs and risk categorisation, improve data discovery and consider setting up database journals. In essence, don't just  tweak the current model. The elephant in the room is the real reason data sharing isn't happening as quickly as we might like. No one wants to, whether they are researchers, institutions or companies.