{"id":106350,"date":"2021-10-27T07:00:13","date_gmt":"2021-10-27T11:00:13","guid":{"rendered":"http:\/\/kendallharmon.net\/?p=106350"},"modified":"2021-10-27T17:33:04","modified_gmt":"2021-10-27T21:33:04","slug":"nature-giant-free-index-to-worlds-research-papers-released-online","status":"publish","type":"post","link":"https:\/\/kendallharmon.net\/?p=106350","title":{"rendered":"(Nature) Giant, free index to world\u2019s research papers released online"},"content":{"rendered":"<p>In a project that could unlock the world\u2019s research papers for easier computerized analysis, an American technologist has released online a <a href=\"https:\/\/archive.org\/details\/GeneralIndex\" data-track=\"click\" data-label=\"https:\/\/archive.org\/details\/GeneralIndex\" data-track-category=\"body text link\">gigantic index of the words and short phrases<\/a> contained in more than 100 million journal articles \u2014 including many paywalled papers.<\/p>\n<p>The catalogue, which was released on 7 October and is free to use, holds tables of more than 355 billion words and sentence fragments listed next to the articles in which they appear. It is an effort to help scientists use software to glean insights from published work even if they have no legal access to the underlying papers, says its creator, Carl Malamud. He released the files under the auspices of Public Resource, a non-profit corporation in Sebastopol, California that he founded.<\/p>\n<p>Malamud says that because his index doesn\u2019t contain the full text of articles, but only sentence snippets up to five words long, releasing it does not breach publishers&#8217; copyright restrictions on the re-use of paywalled articles. However, one legal expert says that publishers might question the legality of how Malamud created the index in the first place.<\/p>\n<p>Some researchers who have had early access to the index say it\u2019s a major development in helping them to search the literature with software \u2014 a procedure known as text mining. Gitanjali Yadav, a computational biologist at the University of Cambridge, UK, who studies volatile organic compounds emitted by plants, says she aims to comb through Malamud\u2019s index to produce analyses of the plant chemicals described in the world\u2019s research papers. \u201cThere is no way for me \u2014 or anyone else \u2014 to experimentally analyse or measure the chemical fingerprint of each and every plant species on Earth. Much of the information we seek already exists, in published literature,\u201d she says. But researchers are restricted by lack of access to many papers, Yadav adds.<\/p>\n<p><a href=\"https:\/\/www.nature.com\/articles\/d41586-021-02895-8\">Read it all<\/a>.<\/p>\n<blockquote class=\"twitter-tweet\">\n<p lang=\"en\" dir=\"ltr\">A big development for text-mining? <a href=\"https:\/\/twitter.com\/carlmalamud?ref_src=twsrc%5Etfw\">@carlmalamud<\/a> has released online a giant, free index of 355 billion words and sentences in 107 million science papers. By <a href=\"https:\/\/twitter.com\/HollyElse?ref_src=twsrc%5Etfw\">@hollyelse<\/a> in <a href=\"https:\/\/twitter.com\/Nature?ref_src=twsrc%5Etfw\">@nature<\/a> <a href=\"https:\/\/t.co\/FDE1mlfwbg\">https:\/\/t.co\/FDE1mlfwbg<\/a><\/p>\n<p>&mdash; Richard Van Noorden (@Richvn) <a href=\"https:\/\/twitter.com\/Richvn\/status\/1453042546041573400?ref_src=twsrc%5Etfw\">October 26, 2021<\/a><\/p><\/blockquote>\n<p> <script async src=\"https:\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>In a project that could unlock the world\u2019s research papers for easier computerized analysis, an American technologist has released online a gigantic index of the words and short phrases contained in more than 100 million journal articles \u2014 including many<span class=\"ellipsis\">&hellip;<\/span><\/p>\n<div class=\"read-more\"><a href=\"https:\/\/kendallharmon.net\/?p=106350\">Read more &#8250;<\/a><\/div>\n<p><!-- end of .read-more --><\/p>\n","protected":false},"author":794,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[94,95],"tags":[],"class_list":["post-106350","post","type-post","status-publish","format-standard","hentry","category-blogging-the-internet","category-science-technology"],"_links":{"self":[{"href":"https:\/\/kendallharmon.net\/index.php?rest_route=\/wp\/v2\/posts\/106350","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/kendallharmon.net\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/kendallharmon.net\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/kendallharmon.net\/index.php?rest_route=\/wp\/v2\/users\/794"}],"replies":[{"embeddable":true,"href":"https:\/\/kendallharmon.net\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=106350"}],"version-history":[{"count":3,"href":"https:\/\/kendallharmon.net\/index.php?rest_route=\/wp\/v2\/posts\/106350\/revisions"}],"predecessor-version":[{"id":106354,"href":"https:\/\/kendallharmon.net\/index.php?rest_route=\/wp\/v2\/posts\/106350\/revisions\/106354"}],"wp:attachment":[{"href":"https:\/\/kendallharmon.net\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=106350"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/kendallharmon.net\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=106350"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/kendallharmon.net\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=106350"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}