Showing posts with label workflow management system. Show all posts
Showing posts with label workflow management system. Show all posts

Thursday, 18 October 2018

Getting From Word to JATS XML

In this blog Bill Kasdorf, Principal, Kasdorf & Associates, LLC talks us through a perennial problem and the different approaches to addressing this:


It is a truth universally acknowledged that journal articles need to be in JATS XML but they’re almost always authored in Microsoft Word.

This is not news to anybody reading this. This has been an issue since before JATS existed. Good workflows want XML. So for decades (yes, plural) publishers have been trying to get well structured XML from authors’ manuscripts without having to strip them down to plain text and tag them by hand. (This still happens. I’m not going to include that in my list of strategies because nobody thinks that’s a good idea anymore.)

There are four basic strategies for accomplishing this:
• Dedicated, validating XML editors.
• Editors that emulate or alter MS Word.
• Use Word as-is, converting styles to XML.
• Editors that use Word as-is, with plug-ins.
Here are the pros and cons of these four approaches.


Dedicated, Validating XML Editors


This is the “make the authors do it your way” method. The authors are authoring XML from the get-go. And not just any XML. Not even just any JATS (or whatever XML model). Exactly the specification of JATS that the publisher needs, conforming in every way to the publisher’s or journal’s style guide and technical requirements. This strategy works in controlled authoring situations like the people developing technical documentation. (They’re probably authoring DITA, not JATS.) They’re typically employees of the publisher, and the document structures are exactly the same every day those employees show up to work.

I have never seen this strategy successfully employed in a traditional publishing context, although I have seen it attempted many times. (If anybody knows of a journal publisher doing this successfully, please comment. I’d like to know about it.) This doesn’t work for journals for two main reasons:
1. Authors hate it. They want Word.
2. They have already written the paper before submitting it to the journal. The horse is out of the barn!


Editors that Emulate or Alter MS Word


This always seems like a promising strategy, and it can work when it’s executed well in the right context. The idea is to either let authors use Word, but make it impossible for them to do things you don’t want them to do (like making a line of body text bold when it should be styled as a heading), either by disabling features in Word like local formatting or by creating a separate application that looks and acts a lot like Word.

I have seen this work in some contexts, but for authoring, I’ve seen it fail more often. The reason is No. 1 above. Despite being a lot like Word, it’s not Word, and authors balk at that. These are often Web-based programs, and authors want to write on a plane or the subway. And there’s always No. 2: most journal articles are written before it’s known which journal is going to publish it.

This strategy can work well, though, after authoring. Copyeditors and production staff can use a structured tool like this more successfully than authors can. We’re seeing these kinds of things proliferate in integrated editorial and production systems like Editoria, developed by the Coko Foundation for the University of California Press, and XEditPro, developed by a vendor, diacriTech.


Use Word As-Is, Converting Styles to XML


This is by far the most common way that Word manuscripts get turned into XML today. A well designed set of paragraph and character styles can be created to express virtually all of the structural components that need to be marked up in JATS for a journal article. This is done with a template, a .dotx file in Word, which, when opened, creates a .docx document with all of the required styles built in. And since modern Word files are XML under the hood, you can work with those files to get the JATS XML you need.

The question is who does the styling, and how well it gets done.

Publishers are sometimes eager to give these templates to their authors so they can either write or, post-authoring, style their manuscripts according to the publisher’s requirements. Good luck with that. The problem is that it’s too easy to do it wrong. Use the wrong style. Use local formatting (see above). Put in other things that need to be cleaned up, like extra spaces and carriage returns. Somebody downstream has to fix these things.

Those people downstream tend to be trained professionals, and it’s usually best just to let them do the styling in the first place. This is how most JATS XML starts out these days: as professionally styled Word files. Many prepress vendors have trained staff take raw Word manuscripts and style them, often augmented by programmatic processing to reduce the manual work. These systems, which the vendors have usually developed in-house, also typically do a “pre-edit,” cleaning up the manuscript of many of those nasty inconsistencies programmatically to save the copyeditor work.

This is also at the heart of what I would consider the best in class of such programs, Inera’s eXtyles. Typically, a person or people on the publisher’s staff are trained to properly style accepted manuscripts; eXtyles provides features that makes this easier to do than just using Word’s Styles menu. Then it goes to town, doing lots of processing of the resulting file based on under-the-hood XML. It’s primarily an editorial tool, not just a convert-to-XML tool.


Use Word As-Is, With Plug-Ins


This is not necessarily the same as the previous category, but there’s an overlap: eXtyles is a plug-in for Word, and the resulting styled Word files can just be opened up in Word without the plug-in by a copyeditor or author. But that approach still depends on somebody having styled the manuscript, and subsequent folks not having messed up the styling. It also presents the copyeditor (and then usually the author, who reviews the copyedits) with a manuscript that doesn’t look like the one the author submitted in the first place.

This tends to make authors suspicious—what else might have been changed?—and suspicious authors are more likely to futz. That’s why in those workflows it’s important to use Tracked Changes, though some authors realize that that can be turned on and off by the copyeditor so as not to track every little punctuation correction that’s non-negotiable anyway.

An approach that I have just recently come to appreciate is what Ictect uses. This approach is not dependent on styles. As much as I’ve been an advocate of styles for years, this is actually a good thing. Styles are the result of human judgment and attention. When done by trained professionals, that’s pretty much okay. But on raw author manuscripts—not.

Ictect uses Artificial Intelligence to derive the XML not from the appearance of the article, which is unreliable, but on the content. Stop and think about that a minute. Whereas authors are sloppy or incompetent in getting the formatting right, they are pretty darn obsessive about getting the content right. That’s their paper.

Speaking of which, in addition to not changing the formatting the author submitted, Ictect doesn’t change the content either. The JATS XML is now embedded in that Word file, but you only see that if you’re using the Ictect software. After processing by Ictect, the document is always a Word document and it is always a JATS document. To an author or a copyeditor it just looks like the original Word file. This inspires trust.

I was initially skeptical about this. But it actually works. Given a publisher’s style requirements and a sufficiently representative set of raw author manuscripts, Ictect can be set up to do a shockingly accurate job of generating JATS from raw author manuscripts. In seconds. Nobody plowing through the manuscripts to style them.

There have been tests done by large STM publishers that have demonstrated that Ictect typically produces fully correct, richly tagged JATS for over half of the raw Word manuscript files submitted by authors, and over 90% of manuscripts can be perfected in less than ten minutes by non-technical staff like production editors. The Ictect software highlights the issues and makes it easy for publishing staff to see what the problem is in the Word file and fix it. That’s because the errors aren’t styling errors, they’re content errors. They have to be fixed no matter what.

In case you think this is simplistic or dumbed-down JATS XML, nope. I’m talking about fully expressed, granular JATS, with its metadata header and all the body markup and even granularly tagged references that enable Crossref and PubMed processing. Not just good-enough JATS. Microsoft Office 365 is not exactly a new kid on the block now, but journal publishers have not made much use of it. As things evolve naturally, more and more authors are going to use Office 365 for peer review, quick editing, corrections and even for full article writing. Since Ictect software creates a richly tagged Word document that can be edited using Office 365, it opens up some interesting workflow automation and collaboration possibilities, especially for large scale publishing.

And if you need consistently styled Word files, no problem. Because you’ve got that rich JATS markup, a styled file can be generated automatically in seconds. For example, in a consistent format for copyediting (I would strongly recommend that), or a format that’s modeled after the final published article format. Authors also really like to see that at an early stage. It’s an unavoidable psychological truism that when an author sees an article in published form she notices things she hadn’t noticed in her manuscript. So you can do both: return the manuscript in its original form, and provide a PDF from the styled Word file to emulate the final layout.

All of the methods I’ve discussed in this blog have a place in the ecosystem, in the right context. I haven’t mentioned a product that I wouldn’t recommend in the right situation. For example, you might initially view Ictect as a competitor of eXtyles and those home-grown programs the prepress vendors use. It’s not. It belongs upstream of them. It’s a way to get really well tagged JATS from raw author manuscripts to facilitate the use of editorial tools, without requiring manual styling. It’s the beginning of an Intelligent Content Workflow. It’s a very interesting development.

Bill Kasdorf is Principal of Kasdorf & Associates, LLC, a consultancy specializing in accessibility, XML/HTML/EPUB modeling, editorial and production workflows, and standards alignment. He is a founding partner of Publishing Technology Partners 

Website: https://pubtechpartners.com/

Twitter: @BillKasdorf


To find out further information on Ictect visit: http://www.ictect.com/ 

or register for one of their free monthly webinars at: http://www.ictect.com/journal-webinars




Thursday, 22 February 2018

Highlights from the 2018 University Press Redux Conference: A Sponsor's Perspective


Virtusales was delighted to sponsor the 2018 University Press Redux Conference, which was held at The British Library and completely sold out with an atmosphere that reflected it. Filled with lively discussions on policy, open access, disruptive innovation and the opportunities it presents, the importance of publishers providing genuine added value, and even Brexit and Trump.



In the opening keynote, Timothy Wright, CEO at Edinburgh University Press, presented the challenges for 2018 as: monographs, new models in print and distribution, ebooks, content, the skills gaps, open access, institutional support and the financial challenges facing the industry.

Richard Fisher from Yale University Press said that discoverability remains the number one challenge for most university presses and explained how marketing for individual titles has diminished but is still vital. Michael Jubb of Jubb Consulting reiterated this, stating that 50% of university press sales are sold through global retail channels such as Amazon, making it essential for content to be easily discoverable amidst the proliferation in formats, business models and retailers.

As a software supplier operating globally, we were particularly interested in the parallel session on Global, which looked at university presses outside of Europe, UK and USA. Another compelling topic was Digital, which was covered in the plenary session with Allison Belan from Duke University Press and Charles Watkinson from University of Michigan Press, chaired by Nicole Mitchell from University of Washington Press. The session included stimulating debate on the pros and cons of buying vs building systems and platforms. Allison explained that it is of utmost importance for university presses to understand and own their own content, data and business rules, and Duke’s decision to buy in technology expertise and systems allows the press to focus on what they are good at - creating content.

Author engagement and support was another recurring theme throughout the conference with an emphasis put on the need for publishers to add genuine value throughout the supply chain and for authors, contributors and stakeholders to recognise the value added. This was at the heart of Tuesday’s parallel session on Production where delegates heard from Andy Redman from Oxford University Press, Neil Clarke from CPI UK, Bret Freeman from LifeVroom, and Wednesday’s session on Commissioning with Simon Bell from Emerald Group Publishing, Brian Halley from University of Massachusetts Press and Katherine Reeve from Bath Spa University.

On the policy front, the requirement for all long form works to be available as open access in order to be eligible for the 2027 REF was by far the most impacting point mentioned, with Steven Hill from HEFCE suggesting how it could be achieved using different models such as Freemium, author pays, and mission orientated new university presses.

Closing Keynote: Richard Charkin
Finally, Richard Charkin from Bloomsbury Publishing gave an engaging closing keynote on academic publishing. Technology was mentioned in virtually all presentations, in one form or another, and it seems clear that publishers need its support now more than ever. It is essential for platforms and systems to be flexible enough to manage metadata at content, article and chapter level, bring efficiencies to the production process and supply chain, and liberate publishers so that they can focus on creating, curating and enriching content to engage with their consumers.

At Virtusales we are committed to streamlining publishers’ workflows and bring efficiencies to their business processes with innovative software, exceptional service and a collaborative approach. Our recent white paper looks at the current landscape of academic and educational publishing, some of the disruptions that publishers are facing in today’s arena and ways in which they can capitalise on the changing environment.

Virtusales Publishing Solutions is the creator of the Biblio suite of publishing software and works closely with some of the world's leading academic, scholarly, professional and trade publishers including Harvard University Press, Bloomsbury Publishing, Manchester University Press, Pearson Education, Penguin Random House, Hachette and Macmillan Publishers.


Find out more about Virtusales and the Biblio suite or request our white paper, please visit our website: www.virtusales.com
Follow us on Twitter: @virtusales 
LinkedIn: https://www.linkedin.com/company/virtusales-publishing-solutions/

For further information on the University Press Redux conference and to access slides and audio from the event please visit: https://www.alpsp.org/UPRedux

Tuesday, 19 September 2017

The Ability to Capitalize on Timely Research


In this guest blog, Mr. Srinaath Krishnamachari, MD of UI Tech Solutions, draws on his experience of over 17 years in the publishing industry to write about the changing face of content creation. As the Chief Evangelist for technology-driven publishing solutions, he seeks to address the need of publishers to provide relevant customized content separate from freely available information.




In the last five years, publishers have seen readers’ relationship to content change dramatically.  A 2016 Pew Research study stated that a majority of the public—62% in the US—get their news from social media rather than traditional media outlets.  A 2016 BioMed Central study states that physicians are rapidly turning to social media in order to immediately share health and research information with their patients and each other.


RESEARCHERS PUBLISH DIRECTLY

The constant access to breaking news, studies, and research through websites and social media, often for free via social media or open access, has challenged the timing of traditional publishing cycles and threatened the methods by which publishers tend to generate revenue. Researchers are turning to publishing directly themselves to publicly accessed websites, as Nobel Laureate Carol Greider did last year, or via other channels in order to get the information out into the world, taking valuable information out of the hands of traditional publishers.

As publishers try to adapt to this changing face of their audience and the industry, they must focus on trying to release research more quickly in order to respond to time-sensitive issues and creating sophisticated metadata that allows for content to be easily found amid the deluge of content.


PUBLISHERS ADAPTING TO THE SPEED OF RESEARCH

To help publishers adjust to this brave new world, PageMajik has created a product suite based on the personal consulting model their parent company S4Carlisle has offered publisher clients for nearly two decades.  PageMajik allows publishers to automate significant portions of the publishing process from author submission to final production in order to improve efficiency and timeliness of content. Current publishers using the system have noted that their efficiency has improved by an average of 40%.
For publishers and authors publishing time-sensitive research, cutting the publication cycle in half can not only make their research more relevant but also speed up technological advances, scientific discovery, and medical breakthroughs.


SIMPLIFYING THE PROCESS

By simplifying the publishing process and automating some of the more detailed and time-consuming technical work, publishers can focus more directly on the much more important task of identifying notable research.

PageMajik simplifies the process by optimizing and organizing existing content and resources while adapting to a publisher’s current systems and workflow.  Working in a web-based authoring environment and InDesign, PageMajik does not require additional training and allows everyone along the publishing cycle to work on the document.  PageMajik also performs a number of detailed tasks that are time-consuming for publishers, such as identifying inconsistencies and anomalies in usages and forms of words (hyphenations, allowed prefixes, precise usages) using pattern-based rules, and grammatical discrepancies using built-in English language rules.


MAKING RESEARCH EASY TO FIND

A challenge that publishers and researchers alike face upon release of published research is how to reach the audience for the work.  Significant studies have been done on the lag between publication and discovery of research, with academics struggling with how to find the right information amid the deluge of content.

PageMajik has created the ability for publishers to do chapter level metadata tagging, allowing for a deeper level of search functionability and, thus, easier discoverability.

As the audience for research and scholarly publishing expands and changes even more toward the digital and direct outreach, publishers must find the ways to work quickly and easily adapt their current systems to stay not only competitive but viable.  PageMajik provides a simple, cost-effective method for publishers to continue to be a vital part of the publishing and research process.

About PageMajik:  PageMajik is a publishing workflow management system that combines all of the individual steps of the publishing process into a seamless product suite to improve workflow and efficiency.  PageMajik’s product suite has an in-built Content Management System that facilitates storage, retrieval and reuse of data at any given time. The CMS has been customized to suit publishing workflows, with version control features and user access control.

Facebook: https://www.facebook.com/PageMajik-145323369390543/
Twitter: https://twitter.com/WeArePageMajik
LinkedIn: https://www.linkedin.com/company-beta/13394487/

PageMajik is a proud sponsor of the ALPSP 10th Anniversary Annual Conference