Wednesday, 22 January 2014

Laura Cox on institutional identifiers internally and throughout the supply chain

Laura Cox: do a data health check

Laura Cox, Chief Financial and Operating Officer at Ringgold kicked off her talk with a focus on what constitutes healthy data. Good quality, reliable and consistent data helps make good decisions. You gain insight into customers and business relationships as well as support strategic planning, decision making and ongoing business operations.

Poor data has real consequences. It is hard to get a true picture of relationships with institutions, can lead to a lack of quality author (and affiliation) data and an inability to see overlap between authors, members and customers. It can drive inaccurate holdings and revenue reports leading to protracted time and effort which can cost your business time and money. Healthy records are complete, accurate, free of duplications, current, consistent and conform with standard identifiers. 

What are unique identifiers and how can they help? 

They are numeric or alpha-numeric designations which are associated with a single entity. Entities can be institutions, persons or pieces of content. They enable the disambiguation of each entity and provide a proper understanding of customer, author, reader or institution as well as a proper identification of content object, article, product or package. They can also be used internally or in conjunction with external partners.

Why should we worry about data now? 

Cox cited the 2012 STM Report (Ware, M and Mabe, M. The STM Report, 2012) which stated that the number of researchers and the number of article are both increasing by 3% per annum. The number of journals is increasing by 3.5% per annum and growth in China has been in double digits for over 15 years. At the same time there is increased demand for anytime/anywhere access while library budgets are frozen or being cut, less money for more content.

Institutional Identifiers can be used for disambiguation (e.g. which UCL?), consolidating different versions (many ways of describing the University of Oxford and its institutions). They provide a hierarchy view (institute within an institution) and reinforce uniqueness. This means you can use them for a gap analysis. 

The Kafka-esque 'Identifiers identified' slide
The main challenge is around multiple data sources. There are system data silos, multiple locations - geographic data silos, data entered by different people for different purposes, data from 3rd parties in the supply chain and data from bought in sources. These things aren’t integrated. Typical publisher systems include financial, CRM or sales databases, authentication system, fulfilment, usage statistics, submissions systems and so on.

Cox advised that the first thing to do is think about your data and implement a data governance plan. What data is held, where and how is it accessed? How can it be used to benefit business and work across silos? But always bear in mind where are you now and where do you want to go?

Another recommendation was to improve data capture. If you can, use web forms as they minimise variance in data input. Implement required fields, use date validation and at a minimum use naming conventions. There are a number of tools such as address validation, postcode look-up, institution validation/lookup. Avoid free-text fields and make institutional identifiers a requirement.

You can use an institutional identifier as a lynch pin to link internal systems for better data integration. It can prevent duplicate account creations, help keep data up-to-date and systems synchronized. It also enables staff to use data more effectively, break down silos, simplify data transmission and provide more insight and power to analyse and understand the business.

So what can you do now?
  • Engage with the problems
  • Think about resources. Time? Money? Systems?
  • How do you want it to work – look at priorities
  • Have a data governance policy
  • Appoint a data champion and document everything
  • Create some basic rules for data entry
  • Use universal identifiers to clean and link your data
  • Work with suppliers and customers to use institutional identifiers ot strengthen the supply chain.

Data, the universe and everything. How data can drive your business

Melinda Kenneway: loves data
Melinda Kenneway, Chair of Data, the universe and everything: How data can drive your business seminar, opened the day by emphasizing how data underpins modern business. She summed up the premise for the day by quoting Jeff Weiner from LinkedIn:

‘Data really powers everything that we do.’ 

If we’re not selling content, what are we selling? Data is also becoming a product or service itself now. An era of big data can help us understand and deliver a much better experience for customers In 2012, 90% of all the data that existed in our entire history has been created in the previous 2 years. We need to get the basics of data right, but also have a strategy.

It’s knowing what you need to know rather than trying to find out about everything. Have clear objectives and stick to them. Don’t forget to consider privacy and legal issues about data. Deliver useful services that build on what the market thinks is acceptable. And if you don’t like data now? You should get out of marketing.

Colin Meddings from DataSalon focused on why you should care about cleaning up your data. To start with it can embarrass or annoy your customers when it is wrong and it’s open to error when users don’t put in info correctly.


Colin Meddings: not vague about data
Continuing with quotes from the great and the good, he cited an observation from 2013’s UKSG conference by Liam Earney from JISC:

‘Publishers can’t even tell a library what they subscribe to.’ 

We’re too vague about absolute fundamental data. The basic process of publishing involves: submit, review, copy edit, publish, purchase, read, cite, licence. All of these aspects generate data. The volume of data is huge and all these areas are where errors or mistakes can happen and where data can conflict.

Reflect on your efficiency with dealing in data. Do you struggle with data wrangling where you have to continually review and refresh datasets? See data as an asset in your business. It is worth investing in. You will get a return. Don’t forget that there’s a legal requirement around data to ensure it is accurate and fit for purpose. He recommended the DQM Data Governance Maturity Model. Be aware. Become reactive. Turn this into proactive. Then make it a managed process. And finally, it becomes optimal. 

Meddings outlined the seven deadly sins of data quality:

  1. Missing data 
  2. ‘Siloed’ data 
  3. Invalid data 
  4. Out of date info (people move around) 
  5. Inconsistencies 
  6. Right information, wrong field 
  7. Duplicate and conflicting data. 

Put some effort and resource into sorting data. Task to someone specifically, create champions, set targets. It’s great if you can employ a team, but if you can’t, get senior management buy-in and work with those who have aptitude/passion for data. Start with one problem, don’t try to fix all at once.

Do a data audit. It can be manual. People who work with data everyday will know where the problems are. Sometimes an automated audit is better – good at finding information about your data for identifying empty fields, etc.

Wednesday, 15 January 2014

Kurt Paulus on ALPSP International Conference 2013: Part 1 - setting the scene

Setting the scene at the ALPSP conference
This is the first in a series of reflections on the 2013 ALPSP International Conference by Kurt Paulus, former Operations Director at the Institute of Physics, and long time supporter of ALPSP. Our thanks go to Kurt for capturing the sessions. If this whets your appetite, save the date for the 2014 conference, 10-12 September, Park Inn Heathrow London.

"Page fright: where to begin? Six plenaries and six parallel sessions in this sixth ALPSP International Conference, with six x six papers presented in all. How to make sense of all this in a number of screens small enough to entice anyone to venture beyond screen one? We shall see. Suffice to say that the 250 or so registrants had a varied, instructive and enjoyable time in the big marquee outside, and other facilities inside The Belfry near Birmingham and each will have carried away new insights, contacts and friendships, repaying the three days spent away from the office.

Inevitably this account of the conference is but a sketch. Detailed presentations can be viewed on the ALPSP website and YouTube channel. Early accounts were posted on the ALPSP blog.

Unsurprisingly this conference was all about change – technical change, changing customer profiles, changing participants in the great scholarly publishing endeavour, changing needs. Nothing new then, as the predecessor conferences and seminars have also been about change, albeit at a slower pace, and change and uncertainty will continue to be with us. One reason for the success of these conferences is that they allow us to take comfort from the support of our fellow publishing professionals and their willingness to share their experiences with us.

Setting the scene

‘Waving – or drowning?’ was how Tim Brooks, CEO of BMJ, headed his keynote talk opening the conference, adapting the title from a poetry collection by Stevie Smith. From his experience of the newspaper industry and his membership of the Cabinet Office’s digital advisory board he was able to draw lessons from other fields.

Waving - or drowning? asked Tim Brooks.
Other Belfry visitors may agree.

Modern life is very complex and change is not always predictable.

While leadership may have a vested interest in stability and the status quo (cultural obstacles to change), change requires agile responses: multi-level, high-speed, self-correcting.




“Doubt is not a pleasant state, but certainty is a ridiculous one” Voltaire

The newspaper industry has been and still is responding to the digital revolution and different publishers are coming up with different business models: the traditional, paper-based one is still durable but not eternal, and it is not yet entirely clear whether pay walls (The Times) or free access (The Guardian) will be the best approach. Nor is it clear which of the available structural options will be most viable. Even the metrics can fool you: new services will not outperform established ones on old-service metrics like profit, at least initially.

Some “surfing tips for the digital breakers”:
  • Ensure those who know, have a voice – the Inuit have known about climate change since the 1960s but nobody listened.
  • Get all publishing functions involved and enthused and get them to imagine the future.
  • Get people in with outside experience.
  • Treat staff as volunteers: explain what they do and why they do it and celebrate outcomes. Look after the team and yourself!
  • Keep talking and listening, especially when things go wrong: the good news culture is unhelpful.
  • Honour the rule of five: Five positives against one negatives is a good indicator of success.
Some of these tips recurred in other sessions."

Rapporteur
Kurt Paulus, Bradford-on-Avon

Monday, 16 December 2013

Colin Meddings: Why data quality matters.

Colin Meddings is the Client Director at DataSalon. Colin will be one of the speakers at the forthcoming ALPSP seminar Data, the universe and everything taking place in January.

Here, in a guest post, he reflects on why good quality customer and internal data is important for scholarly publishers.


'Only four types of organisations need to worry about data quality: Those that care about their customers; Those that care about profit and loss; Those that care about their employees; and Those that care about their futures.' – Thomas C. Redman (2006)

Over recent years publishers have had to overcome many hurdles in the digital world, such as making content available online, managing complex consortia deals, creating new packages of content and tracking usage statistics. The result of all this digital activity is vast amounts of data. However, the pace of change can often distract from the careful governance of this data, leading to gaps, inconsistencies and inaccuracies.

But why does the quality of all this data matter so much? Good data is your most valuable asset, and bad data can seriously harm your business and credibility…

What have you missed? 
At a management level, poor data quality equates directly to poor visibility of key trends in the growth or decline of certain products or markets. At the contact level, you may miss out on valuable sales opportunities if email address fields aren’t filled out correctly or customer names are wrong. Having good data will help deliver better customer service and enhance your reputation, and it means you can make better selections for targeted prospecting, cross-selling and up-selling.

When things go wrong.
Bad data can lead to ‘accidents’ and wrong decisions or actions which can affect customer confidence. You’ve spent time building up a valuable customer list – so it’s important not to waste this by sending campaigns to the wrong people, or with messages which don’t match their interests, or to out-of-date or deceased contacts. Data quality issues can also cost you money directly – for example if invoices or renewal notices are sent to the wrong recipient, or at the wrong time.

Making confident decisions. 
Data quality matters most of all because it enables your staff and management team to really trust the accuracy of the reports and analysis they’re given. Without that confidence, apparent trends or new opportunities will always leave you wondering whether they really present a true picture. But with a complete and accurate view of your customers and prospects, comes the confidence to make well informed business decisions and commit fully to your strategic planning.

So, data quality is a very important foundation for a publisher’s entire business planning process and customer contact strategy. Good data quality will allow your business and its reputation to grow and flourish.

Data quality is just one of the topics in the forthcoming ALPSP seminar Data, the universe and everything. Other areas covered will include the use of institutional and personal identifiers in the scholarly publishing supply chain, publisher metadata, data relating to open access publishing and some case studies from publishers who have tackled data issues.

This post originally appeared on DataSalon’s own blog From the Armchair.

Wednesday, 4 December 2013

Frank Stein on Watson and the Journey to Cognitive Computing

Frank Stein on cognitive computing
Frank Stein from IBM outlined their project Watson and the Journey to Cognitive Computing at the STM Innovations seminar. Data is exploding driven by unstructured data (in descending order: video, image, audio, text, structured data). How do we build a system that can take all this info and build something useful for researchers, doctors, etc?

The Watson and Jeopardy! example shows how they have developed a programme that can match deeper evidence and use temporal reasoning, statistical paraphrasing and geospatial reasoning. The evidence is still not 100% certain, but it is about about likelihood and confidence.

What they learned in Jeopardy
The DeepQA approach can accurately answer single sentence queries with confidence and speed. It is highly dependent on content, content quality, and content formats. They need a combination of technologies to get satisfactory performance (semantic technology, machine learning, information retrieval/search technology, databases and high performance computing techniques). Both structured and unstructured content need to be combined for best results. They now need to extend Watson to handle richer interactions and continuous training/learning.

Here's the IBM video about Watson and the game show Jeopardy!


Watson Decision Advisor in medicine
A data-rich, societally important field helping Watson change how medicine is:
IBM used to produce typewriters
When Stein started, IBM produced typewriters. Now they have 10,000+ products. Their sales agents need help. IBM is building out a portfolio of Watson Solutions including Watson Engagement Advisor for use in situations in which you need stronger ties with constituents and better automated or agent-facilitated conversations. Examples include: bank outreach to customers for cross-sell, cable operator services and support, tax agency advice, etc. 

What's next - Cognitive Computing
Watson is ushering in a new era of computing. We have transitioned from the tabulating systems era to programmable systems era. Now we are moving into a world called cognitive systems era. This is a key technology for a new era of computing that takes into account:
  • Content and learning
  • Visual analytics and interaction
  • Data centric systems
  • Cognitive architecture
  • Atomic and nano-scale.

Sayeed Choudhury reflects on the research data revolution

Sayeed Choudhury
Sayeed Choudhury, Associate Dean for Research Data Management, Johns Hopkins University, kicked off the STM Innovations seminar reflecting on 'The Research Data Revolution'.

There is a new economy of sources of data. The challenge as publishers is to develop services.

Data Conservancy is a community that develops solutions for data preservation and sharing to promote cross-disciplinary re-use. It is about preservation - collect and take care of research data; sharing - reveal data's potential and possibilities; and discovery - promote re-use and new combinations.

Is data different?
Data is the new oil (stated in Qatar, European Commission, etc). McKinsey claimed that data is 4th factor of production and estimates a potential $3 trillion of economic value across seven sectors within the US alone. Todd Park estimates location sensitive apps generate $90 billion of value annually. Policy movements reflect its importance: the White House Office of Science & Technology Policy Executive memorandum and White House Open Government Initiative are two key initiatives.

Collections
Data are a new form of collections though they are fundamentally different in nature. They are created or converted to digital format for processing by machines. Entirely new methods are required to deal with them. They are, in effect, a new form of special collections.

What is 'Big Data'?
There are definitions based on the V's of Big Data (e.g. volume, velocity, variety). What is clear is that it's different from 'spreadsheet science' (or long-tail science). For Choudhury, if a community's ability to deal with data is overwhelmed, it is 'Big Data' - and it's more about 'M's' (methods of lack thereof) than 'V's'.

Services
There's a core of services that span across data from different disciplines and contexts. Archiving is a good example. However, if data collections are basically open, libraries may need to differentiate themselves by the services they offer. They should provide a combination of machine and human mediated services. There will be a set of services that only 'experts' will be able to offer.

Data management layers: curation, preservation, archiving, storage

Understanding infrastructure
Data will require fundamentally new systems and infrastructure. Institutional repositories can be useful gateways, but are not long-term solutions (particularly for 'Big Data'). Libraries will need to operate at scale through an integrated, ecosystem approach to infrastructure. Customised 'human mediated' services are most effective as an interpretative layer on machine based services.

What about publishers?
No one can claim a specific role or act with a sense of entitlement when it comes to data (whether publishers or librarians). The future of data curation is a competition between information graphs. 'Publishing is about content, not format.' - Wendy Queen, Associate Director of Project Muse, Johns Hopkins University Press

Monday, 2 December 2013

International Publishers Association Call for Nominations: 2014 IPA Freedom to Publish Prize

The closing date for nominations for the 2014 IPA Freedom to Publish Prize is 6 January 2014.

The Prize will be awarded on 27 March 2014, during the IPA Congress in Bangkok, and the recipient will receive CHF20,000, thanks to the generous sponsorship of the following publishers: Albert Bonniers Förlag, Elsevier, HarperCollins, Kodansha, Macmillan, OUP, Penguin Random House, and Simon & Schuster.

Nominees can either be publishers who have recently published controversial works in the face of pressure, threats, intimidation or harassment from government or other authorities; or publishers with a long and distinguished history of upholding the values of freedom to publish and freedom of expression.

IPA member organisations, members of the IPA Freedom to Publish Committee, individual publishers, and international professional and non-government organisations working in the field of freedom of expression can nominate candidates for the IPA Freedom to Publish Prize.

Those nominating must explain the reasons behind their choice of candidate in writing (in English, French or Spanish) using the attached form as a template. Nominations should be submitted to the IPA’s Policy Director, José Borghino (borghino@internationalpublishers.org) no later than close-of-business (Geneva time) on 6 January 2014.

More about the 30th IPA Congress and the IPA Freedom to Publish Prize Ceremony: 

The 30th IPA Congress will be held in Bangkok, Thailand, on 25-27 March 2014, and will be hosted by the Publishers and Booksellers Association of Thailand (PUBAT) under the auspices of HRH Princess Maha Chakri Sirindhorn. To see the Program, go to the Congress website.

On the eve of the Bangkok Book Fair (28 March to 8 April), hundreds of publishers from all over the world will participate in the Congress, together with authors, copyright specialists, librarians and officials from around 50 countries and international organisations.

The 2014 IPA Freedom to Publish Prize will be awarded during the Congress on 27 March 2014. Aung San Suu Kyi has been invited to give the keynote speech and award the Prize.

Earlybird online registration for the Congress is available from the Congress website.