{"id":70,"date":"2010-02-10T22:01:54","date_gmt":"2010-02-11T06:01:54","guid":{"rendered":"http:\/\/www.davidclausen.org\/tower\/?p=70"},"modified":"2010-02-10T22:01:54","modified_gmt":"2010-02-11T06:01:54","slug":"web-scale-distributional-similarity-and-entity-set-expansion","status":"publish","type":"post","link":"https:\/\/www.davidclausen.org\/tower\/2010\/02\/10\/web-scale-distributional-similarity-and-entity-set-expansion\/","title":{"rendered":"Web-Scale Distributional Similarity and Entity Set Expansion"},"content":{"rendered":"<p>Web-Scale Distributional Similarity and Entity Set Expansion &#8211;\u00a0Patrick Pantel, Eric Crestan, Arkady Borkovsky, Ana-Maria Popescu and Vishnu Vyas. 2009<\/p>\n<p>This paper implements web scale word similarity metrics and uses this to test performance on a set expansion task.\u00a0 The authors implement a distributed algorithm to compute cosine similarity between context vectors for 500 million terms found in a 200 billion word web dump.\u00a0 Context vectors consisted of pmi between the word and the chunk immediately to the left or to the right although the details of the features are unclear.\u00a0 Using 200 quad core hadoop instances the entire similarity matrix was computed in 5 hours.\u00a0 The authors took advantage of the sparsity of the matrices to build an inverted index.\u00a0 They then decomposed the similarity computation into 3 parts that involved only features in one word, only features in the other word and nonzero features for both words.\u00a0 The inverted index allowed practical computation of the score part relying on both words while the other score parts that relied on individual words could be cashed.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Web-Scale Distributional Similarity and Entity Set Expansion &#8211;\u00a0Patrick Pantel, Eric Crestan, Arkady Borkovsky, Ana-Maria Popescu and Vishnu Vyas. 2009 This paper implements web scale word similarity metrics and uses this to test performance on a set expansion task.\u00a0 The authors implement a distributed algorithm to compute cosine similarity between context vectors for 500 million terms [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-70","post","type-post","status-publish","format-standard","hentry","category-papers"],"_links":{"self":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts\/70","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/comments?post=70"}],"version-history":[{"count":2,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts\/70\/revisions"}],"predecessor-version":[{"id":72,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts\/70\/revisions\/72"}],"wp:attachment":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/media?parent=70"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/categories?post=70"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/tags?post=70"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}