{"id":67,"date":"2010-02-10T21:59:46","date_gmt":"2010-02-11T05:59:46","guid":{"rendered":"http:\/\/www.davidclausen.org\/tower\/?p=67"},"modified":"2010-02-10T22:00:01","modified_gmt":"2010-02-11T06:00:01","slug":"entity-extraction-via-ensemble-semantics","status":"publish","type":"post","link":"https:\/\/www.davidclausen.org\/tower\/2010\/02\/10\/entity-extraction-via-ensemble-semantics\/","title":{"rendered":"Entity Extraction via Ensemble Semantics"},"content":{"rendered":"<p>Entity Extraction via Ensemble Semantics &#8211; Pennacchiotti &amp; Patel (2009)<\/p>\n<p>This paper proposes a new framework for information extraction called Ensemble Semantics.\u00a0 The Authors describe a Knowledge Extraction framework which collects information from multiple knowledge sources.\u00a0 They then use multiple knowledge extractors and feature extractors to extract candidate relations and features of relation reliability.\u00a0 They then rank candidate relations producing the final knowledge relations.\u00a0 The particular system drew on web query logs, web corpus data (600 million docs), structured web data (web tables) and Wikipedia features.\u00a0 They demonstrate combining multiple knowledge extractors (distributional and pattern based) along with features from all knowledge sources significantly improve MAP for the categories of musicians, actors and athletes.\u00a0 Wikipedia features alone significantly improved performance over the baseline system.\u00a0 Including web querry logs and web corpus data further improved performance and subsumed the benefits of using Wikipedia features.\u00a0 The confidence of the knowledge extractors were the most important features.\u00a0 Features extracted from web data were also very important.\u00a0 This data was drawn from query logs, structured table data, and free form web data.\u00a0 Features determining the well formed nature of the Term, popularity of the term and common co occurrence terms were all very useful. The weighting of features varied based on knowledge source illustrating a key advantage of the Esemble Semantics framework.\u00a0 Combining different knowledge sources and using different features to integrate them significantly improves knowledge extraction performance.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Entity Extraction via Ensemble Semantics &#8211; Pennacchiotti &amp; Patel (2009) This paper proposes a new framework for information extraction called Ensemble Semantics.\u00a0 The Authors describe a Knowledge Extraction framework which collects information from multiple knowledge sources.\u00a0 They then use multiple knowledge extractors and feature extractors to extract candidate relations and features of relation reliability.\u00a0 They [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-67","post","type-post","status-publish","format-standard","hentry","category-papers"],"_links":{"self":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts\/67","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/comments?post=67"}],"version-history":[{"count":2,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts\/67\/revisions"}],"predecessor-version":[{"id":69,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/posts\/67\/revisions\/69"}],"wp:attachment":[{"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/media?parent=67"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/categories?post=67"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.davidclausen.org\/tower\/wp-json\/wp\/v2\/tags?post=67"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}