Showing posts with label reproducibility. Show all posts
Showing posts with label reproducibility. Show all posts

Tuesday, 20 August 2019

Spotlight on Ripeta - shortlisted for the 2019 ALPSP Awards for Innovation in Publishing

On 12 September we will be announcing the winners of the 2019 ALPSP 
Awards for Innovation in Publishing.  In this series of posts leading up to the Awards ceremony we meet our four finalists and get to know a bit more about them.  
logo Ripeta
First question has to be: Tell us about your company

Ripeta was founded by a team of three science-related experts who joined forces whilst working at Washington University in St. Louis. Leslie, Anthony and Cynthia, like many of their colleagues, were increasingly burdened with the task of improving science without having the resources to do so. They found that much of the curated data they produced at a data center were used but was not cited. 


They quickly became frustrated and saw the need for a tool and practices to remedy the situation, and the idea for Ripeta was born. The Ripeta software was developed to assess, design, and disseminate practices and measures to improve the reproducibility of science with minimal burden on scientists, starting with the biomedical sciences. In 2017, Ripeta launched its first alpha product.

What is the project that you submitted for the Awards?

Main features and functions of Ripeta

Ripeta is designed for publishers, funders, and researchers. We provide a suite of tools and services to rapidly screen and assess manuscripts for the proper reporting of scientific method components. These tools leverage sophisticated machine-learning and natural language processing algorithms to extract key reproducibility elements from research articles. This, in turn, shortens and improves the publication process while making the methods easily discoverable for future reuse.

Ripeta Software:  Our software is built on a reproducibility framework that includes over 100 unique variables grouped into five categories highlighting important scientific elements and creating a report: 
  1. Study Overview
  2. Data Collection
  3. Analytics
  4. Software and Code
  5. Supporting Material (new)
This report is generated in seconds providing immediate feedback.

image ripeta report


















Portfolio Analysis:   While many publishers have adopted checklists to evaluate reproducibility reporting criteria, the standard practice is for editors or reviewers to manually assess document adherence to guidelines. This review process is neither quick nor consistent. Ripeta offers a critical addition to the scientific publishing pipeline while significantly reducing the time to review and assess manuscripts.   We look across a group of articles based on your criteria and provide an overview, comparisons, and suggested improvements. Do you want to know how your journal or organization is performing overall? Want to know the reporting practices of your grantees? These reports provide insights into reproducibility practices.
graphic Portfolio analysis

Tell us about the team 

Dr. Leslie McIntosh, PhD is the founder and CEO. She has led multi-million dollar projects building software and services for data sharing and reuse. Leslie has a focus on assessing and improving the full research cycle and making the research process reproducible.


Anthony Juehne, MPH is the Chief Science Officer with speciality skills in epidemiology and biostatistics. His current work focuses on developing best-practices for conducting and reporting clinical research to enhance reproducibility, transparency, and accessibility.



Cynthia R. Hudson Vitale, MLIS is the Chief Information Scientist. She has worked with faculty on projects to facilitate data sharing and interoperability while meeting faculty research data needs throughout the research life-cycle. Her current research seeks to improve research reproducibility, addressing both technical and cultural barriers.

In what ways do you think Ripeta demonstrates innovation?

Research reproducibility is increasingly important in the scholarly communication world, yet researchers, publishers, and funders do not have a streamlined method for assessing the quality and completeness of the scientific research.

Ripeta aims to make better science easier by identifying and highlighting the important parts of research that should be transparently presented in a manuscript and other materials.  By detecting and predicting reproducibility in scientific research we provide a “credit report” for scientific publications. Our aim is to improve science and ensure resources are well spent, and we offer a pre peer-review to a paper and report post-publication.

The Ripeta solution is unique because we have identified the components across scientific fields necessary to responsibly report a scientific process. We are using machine learning and natural language processing to programmatically extract these components from scientific manuscripts and present them in a user-friendly report.

Ripeta helps save time, money, and improve reputations.

For publishers, Ripeta offers a quick assessment of a submitted manuscript. It’s hard to find reviewers and the reviewers are experts in specific areas. Ripeta allows a rapid check of important elements that should be in the manuscript and presents a report for editors and reviewers. This is measured in time to complete a pre-review.

For researchers, Ripeta offers a means to improve the manuscript. By rapidly reviewing the manuscript, the report will highlight which elements are missing and which elements have been found in a machine-readable manner. This allows researchers the opportunity to improve their manuscript before submission. This is measured through Ripeta report completeness from first to final submission.

What are your plans for the future?

The Ripeta go-to-market strategy is focused on developing tools for a subscription-model compatible with the needs of publishers and funders where users can assess a single publication. Subscribers will be charged a fee per report with an optional Ripeta software enterprise edition for robust analytics and bulk manuscript evaluation. We are currently engaging and conducting pilots with multiple publishers, universities, and researchers.

Our long-term goals include developing a suite of tools across the broad spectrum of sciences to understand and measure the key standards and limitations for scientific reproducibility across the research lifecycle and enable an automated approach to their assessment and dissemination.
  • Enable researchers to upload a single manuscript at a time at no charge;
  • Work with  Pre-print services (e.g., bioRxiv), who could charge to have a ripetaReport linked to the pre-print; and,
  • Work with large research and development firms, who would purchase enterprise installations for private hosting.
website: https://www.ripeta.com/
twitter: https://twitter.com/ripeta1

See the ALPSP Awards for Innovation in Publishing Finalists lightning sessions at our Annual Conference on 11-13 September, where the winners will be announced.

The ALPSP Awards for Innovation in Publishing 2019 are sponsored by MPS Ltd.







Wednesday, 7 August 2019

Spotlight on Scite - shortlisted for the 2019 ALPSP Awards for Innovation in Publishing

On 12 September we will be announcing the winners of this year's ALPSP Awards for Innovation in Publishing.  In this series of posts, we meet the finalists to learn a little more about each of them.

In this post, we hear from Josh Nicholson, co-founder and CEO of scite.ai

Tell us a little about your company

logo scite
The idea behind scite was first discussed nearly five years ago in response to a paper from Amgen reporting that out of 53 major cancer studies this company tried to validate, they could only successfully reproduce 6 (11%). This paper sparked widespread media coverage and concern and since then this problem has come to be known as the “reproducibility crisis.” While this paper received the most attention, perhaps because the numbers are so dire, it was not the first or only paper to reveal this problem. Indeed Bayer had reported similar findings in other areas of biomedical research, while non-profit reproducibility initiatives revealed the problem in psychology and other fields, suggesting a systemic issue. This is worrisome, to say the least, because scientific research informs nearly all aspects of our lives, from how you raise your children to the drugs being developed for fatal diseases and if most work is not strong enough to be independently reproduced, we are wasting billions of dollars and impacting millions of lives. scite wants to fix this problem by introducing a system that identifies and promotes reliable research.

We do this by ingesting and analyzing millions of scientific articles, extracting the citation context, and then applying our deep learning models to identify citations as supporting, contradicting, or simply mentioning. In short, allowing anyone to see if a scientific article has been supported or contradicted.

As a funny, aside my co-founder, Yuri Lazebnik, and I first proposed that someone else, like Thomson Reuters, Elsevier, or NCBI should implement the approach used now by scite. After some waiting, we realized that if we wanted it to exist we would need to build it ourselves and here we are, five years later, with over 300M classified citations citing over 20M articles!


Tell us a little about how it works and the team behind it

As mentioned, scite is a tool that allows anyone to see if a scientific paper has been supported or contradicted by using a deep learning model to perform citation analysis at scale. In order to do this, we need to first extract citation statements from full-text scientific articles, which in most cases means extracting citation statements out of PDFs. To accomplish this, scite relies upon 11 different machine learning models with 20 to 30 features each. This is very challenging as there are thousands of citation styles and PDFs come in a variety of different formats and quality. We’re fortunate to have Patrice Lopez on the team, who has been developing the tool to accomplish this for over ten years. Once we’ve extracted the citations from the articles, we use a deep learning model to classify citations as supporting contradicting or mentioning.

To show the utility of the tool, I like to show my PhD research as it is seen with and without the lens of scite. This study looked at the effects of aneuploidy on chromosome mis-segregation, that is, if you add an extra chromosome to a cell does it make more mistakes during cell division. Our work was published in eLife and was a collaboration between our lab at Virginia Tech, and labs in Portugal, and at the NIH. It has been cited 40 times to date and viewed roughly 4,000 times. In general, these features are what we as a community look at when assessing a paper–who the authors are, the prestige of the journal it appears in, affiliations, and some metrics like citations and perhaps social media attention (Altmetrics). This information is used to decide if we want to read or cite a paper, if we want to promote this author, join their lab, or give them a grant. These are our proxies of quality. Yet, none of them have anything to do with quality. With scite, in just a few clicks, you can see that my work has been independently supported by another lab (i.e. it has a supporting cite). To me, this is something like a super power for researchers, because without scite one would need to read forty papers to find this information or consult an expert and even then they might miss it.
screenshot demonstrating cite



To make scite happen requires a special team and I believe that is what we have created and continue to create at scite. I like to joke that scite is a multinational corporation with offices in Kentucky, Brooklyn, France, Germany, and Connecticut. While true, it is not an entirely accurate representation of the company, just as citation numbers are not an entirely accurate representation of a paper. In fact, scite is a small team of six scientists and developers united not by geography but by a passion to make science more reliable. 

In what ways do you think it demonstrates innovation?

The idea behind scite has been discussed as early as the 1920’s, as there exists a similar system in law called Shepardizing (lawyers need to make sure they don’t cite overturned cases  as they will quickly lose their argument this way). However, despite such discussions happening nearly a hundred years ago and multiple attempts to bring something like scite to fruition, even by juggernauts like Elseiver, it did not happen until scite came to life. scite is innovative in that in unlocks a tremendous wealth of information by successfully pushing the latest developments in technology to its limits. With that said, there is so much that we still need to do and we’re excited about the future and working with many stakeholders in the community.

What your plans for the future?

In the near future, we think anywhere there is scholarly metadata, there is an opportunity for scite to provide value. We are working with publishers to display scite badges, citation managers to display citation tallies, submission systems to implement citation screens and are in discussions with various pharmaceutical companies to help improve the efficiency of drug development. Moreover, we will start to expand out citation analytics from articles, to people, to journals, and institutions.

Longer term, we envision scite as being the place where people and machines go to identify reliable research and researchers. We have plans to explore micro-publications, so as to offer more rapid feedback into our system, plans to further invest in machine learning to see if we can predict citation patterns as well as promising therapeutics in drug development, and I think much more that we can’t even predict right now. The scientific corpus is arguably the most important corpus in the world. It’s a shame that it is easier to text mine twitter than it is cancer research. However, it’s also an opportunity, one which we’re seizing now.

photo Josh Nicholson
Josh Nicholson is co-founder and CEO of scite.ai, a deep learning platform that evaluates the reliability of scientific claims by citation analysis. Previously, he was founder and CEO of the Winnower (acquired 2016) and CEO of Authorea (acquired 2018 by Atypon), two companies aimed at improving how scientists publish and collaborate. He holds a PhD in cell biology from Virginia Tech, where his research focused on the effects of aneuploidy on chromosome segregation in cancer.

Websites
https://scite.ai/
Chrome plugin: https://chrome.google.com/webstore/detail/scite/homifejhmckachdikhkgomachelakohh
Firefox plugin:
https://addons.mozilla.org/en-US/firefox/addon/scite/

Twitter:
@sciteai

See the ALPSP Awards for Innovation in Publishing Finalists lightning sessions at the ALPSP Conference on 11-13 September. The winners will be announced at the Dinner on 12 September.

The ALPSP Awards for Innovation in Publishing 2019 are sponsored by MPS Ltd.  

Wednesday, 30 May 2018

Examining Trust and Truth in Scholarly Publishing

In this latest blog, Helen Duriez, from our Professional Development Committee, reflects on how our current webinar series Trust, Truth and Scholarly Publishing webinar series came together.  


Oh, how the world turns. I used to think that Donald Trump running for US president was a fine joke. I used to think there was no way the UK would choose to go it alone when it could be a part of the collective economic might of the European Union. Turns out, the voting public in the US and UK had very different ideas to those of this naïve millennial back in 2016.  

Two years on, it’s become apparent that a large part of the success of these two major political campaigns was their ability to leverage personal belief systems. People are more likely to believe what they read if it aligns with their pre-existing belief system or if it taps into a feeling of existential threat, causing them to disregard evidence to the contrary. Ironically enough, there’s research that backs up this theory, and the concept even has a name – post-truth. You might have heard of it.

Now, what people choose to believe (or not) is tightly interwoven with what we choose to tell them, and how. In scholarly publishing, most of our jobs involve disseminating complex information in one form or another. With research output higher than ever before, there’s a lot of complicated stuff to explain – not just to academics and practitioners, but to the general public as well. Scientists are used to working with ambiguities, although that doesn’t mean they always navigate the rocky terrain of uncertainty safely. And what about lay audiences, who give as much weight to opinions as to facts?

The team at ALPSP felt this topic warranted further exploration, and so a small group of staff and volunteers (I’m one of the latter) have taken it upon ourselves to put together a series of webinars looking at some of the issues and opportunities in scholarly publishing today. Here’s how the series pans out…

Publishing without perishing

In case you missed the first webinar in the Trust, Truth & Scholarly Publishing series, go – sign up and download it. Seriously, do it. Yes, as one of the organisers I may be a little biased, but even knowing what I was about to listen to didn’t stop me from being motivated and left feeling a little awe-inspired as Richard Horton gave us a passionate, powerful reminder of what early journal publishers set out to achieve, and the obligations we still have to society today. Jason Hoyt follows up with some practical thoughts about how publishers can succeed in a post-truth world.

The reproducibility opportunity

In last week's webinar, now available for download too, and highly recommended,  Catriona Fennell, Rachel Tsui and Chris Chambers explored how the concept of reproducible science represents an opportunity, rather than a threat, when it comes to getting to the truth, the whole truth, and nothing but the truth. The traditional journal publishing model doesn’t have much time for replication studies (not original research, don’t ya know) or registered reports (findings, please!), but things are starting to change…

Public engagement with scholarly research

The process of communicating a new piece of scientific research to the world can sometimes feel a little like a game of Chinese whispers. When the description of a complex concept or process is shortened and reworded in order to reach a new audience, it’s meaning can change subtly. I’ve seen more than one twitter spat debating the latest “scientists have found…” health fact, and there are those who have built careers around addressing some of these misrepresentations.

So, what tools can those of us in scholarly communication use to instil trust in our content? In our last webinar we are joined by three industry communicators Tom Griffin, John Eggleton and Eva Emerson to find out.  You can register here for this final webinar.

For more practical information on the series, including how to get a members’ discount, see here.

Helen Duriez is a Product Manager at Wiley, specialising in digital strategy and planning. With over 12 years’ experience in the publishing industry, Helen has previous worked at the Royal Society, Macmillan and OUP. She gets out of bed for open science and avocado toast.



Tuesday, 18 October 2016

Professor Brian Nosek on increased openness and the credibility of science

Brian Nosek from the Center for Open Science at the University if Virginia gave the keynote talk at the 2016 STM Frankfurt conference.

He asked what is it that publishers can do to help scientists be successful? Scientists are constrained by what is human about them, but the reality and our experience of reality are not the same thing. We all have mental modules that want us to see what we want to in order to reinforce our beliefs. So the brain imposes understanding on what it sees.

Applying sociological approaches to the study of science, there are Norms versus Counternorms e.g. communality versus secrecy; universalism versus particularism (evaluate research on own merit or evaluate research by reputation); disinterestedness versus self interested.

There are stark differences when you unpack what researchers believe about how they do science (they do it for the norms) compared to what they observe about themselves (more counter norms creep in) and what they observe of others (almost all counter norms).

The primary challenge is that incentives for individual success are focused on getting it published, not getting it right. The choices that scientists make when analysing the data can impact on the results. And unless you see the source data, you cannot understand why this is. So how do we get researchers to be more transparent and reproducible in their work?

Barriers include perceived norms, motivated reasoning, minimal accountability, and the ubiquitous 'I am busy'. What can be done about it? Look at the rewards that we need and the means to get them. What if you added rewards for transparency and reproducibility in the research process? What if you were to diversify rewards to include data and materials so there is recognition for research content?

Why is this tricky? There are a lot of stakeholders: universities, founders, publishers, societies which creates a complex group of issues. If there are desired behaviours and people don't know they are happening (to increase credibility of their project) you need to raise their profile and boost credibility of research.

At the Center for Open Science they have created badges for this and defined what this means. Badges are symbols and in real life these are powerful, significant indicators of who you are. They are promoting an open research culture using standards, with three data sharing levels:

  1. Article states whether data are available, and if so, where to access them
  2. Data must be posted to a trusted repository. Exceptions must be identified at article submission.
  3. Data must be posted to a trusted repository, and reported analyses will be reproduced independently prior to publication.

So far there are 749 journals from 62 organizations that are signatories of the top level of guidelines. Another focus has been round preregistration of research studies. The Registered Reports workflow is: Design > Collect & Analyse > Report > Publish. Peer review would usually happen between Report and Publish, but they move it back to between Design and Collect. If reviewers don't know what the results are, they are incentivised (along with the researcher) to make it the best study possible. They have 38 journals so far who have committed to making registered reports happen.

They are developing the possibility of partnerships between publishers and funders. A review report is submitted to both and, if acceptable to funder and publisher, it gets the go ahead, providing an efficiency step for all groups.

One of biggest challenges to reproducibility is completely mundane: labs lose materials and data all the time. People have multiple, personal systems of data preservation. What can you do to mitigate this? Adopt the TOP Guidelines. Adopt Badges. Adopt Registered Reports. Partner on preregistration
and partner on open data with OSF.

Brian Nosek is Executive Director of the Center for Open Science and Professor in the Department of Psychology at the University of Virginia. He gave the keynote at the 2016 STM Frankfurt Conference.
SaveSave