Thursday, August 07, 2008

Google Translation Center: The World’s Largest Translation Memory - GigaOM

Google Translation Center: The World’s Largest Translation Memory - GigaOM: "Google has been investing significant resources in a multi-year effort to develop its statistical machine translation technology. Statistical MT works by comparing large numbers of parallel texts that have been translated between languages and from these learns which words and phrases usually map to others — similar to the way humans acquire language. The problem with statistical MT is that it requires a large number of directly translated sentences. These are hard to find, and because of this SMT systems use sources like the proceedings from the European Parliament, United Nations, etc. Which are fine if you’re writing in bureaucrat-speak, but aren’t so great for other texts. Google Translation Center is a straightforward and very clever way to gather a large corpus of parallel texts to train its machine translation systems.

Part machine translator and part translation memory (a sort of search engine for translation that helps translators to recall translations), GTC will help translators by providing a free, global translation memory, and in turn drive costs down by reducing the amount of work needed to complete a text. It will help Google by providing an excellent source of high quality parallel texts that can be fed back into the statistical translation systems."


Tuesday, July 29, 2008

CUIL vs Google

p2pnet news » Blog Archive » CUIL vs Google: "“Rather than assigning priority to pages based on inbound links as Google does (”Pagerank”), Cuil analyzes the content of Web pages to divine their relevance to a search query. Costello bristled when I asked if this was a semantic search engine like PowerSet (recently sold to Microsoft). Costello said Cuil’s search is ‘contextual,’ and that, ‘we’re trying to understand the real world, not the Web’.”
Cuil claims to have better search results than Google and others, “based on how they index websites,” says TechCrunch, going on:
“They do not simply catalog keywords on a site and then rank the site based on its importance. They also work to understand how words are related (France - cheese - wine, for example), to return more relevant results to users. This is a semantic approach to search, but very different from Powerset’s natural language approach.”
Powerset, “uses artificial intelligence to try to understand what sentences on a website actually mean” but Cuil, “simply tries to properly categorize and file a web page, even if the category name doesn’t appear on the site.”"


Thursday, July 24, 2008

Semantic Search Arrives at the Web

Semantic Search Arrives at the Web: "Semantic search has attracted a lot of attention in the past year, largely due to the growth of the semantic web as a whole. The term semantic search itself is popular enough to be considered overused. The term refers to searching large semantic web datasets, which is a typical problem for semantic web search engines such as Swoogle, Sindice, SWSE, Falcon-S, and Watson. The term also refers to methods of searching web documents beyond the syntactic level of matching keywords. This article discusses semantic search in this second sense." (Article by Peter Mika @ Yahoo)


Monday, May 26, 2008

Teragram Integrates Linguistic Tools with Apache Lucene

Teragram Integrates Linguistic Tools with Apache Lucene: "'Lucene is expanding its user base to high-profile corporate and consumer-facing websites around the world, and quickly becoming the open source alternative to traditional enterprise search,' said Dr. Yves Schabes, president and co-founder of Teragram. 'We're happy to provide Lucene users with language processing enhancements so they can meet the high-performance standards of traditional enterprise search engines, while still enjoying the freedom of the open source experience.'"


Monday, May 12, 2008

Powerset, with new search technology, launches today - SiliconValley.com

Powerset, with new search technology, launches today - SiliconValley.com: "Now based in San Francisco, the core of Powerset's technology was developed at Xerox's Palo Alto Resarch Center (PARC), which is famous for incubating breakthroughs like the computer mouse and the graphical user interface. Pell's co-founder Lorenzo Thione is a research scientist who has worked at CommerceNet consortium and the Fuji-Xerox Palo Alto Laboratory. About 25 of Powerset's 60 employees have Ph.Ds., mostly in computational linguistics.
Google has also hired dozens of specialists in computational linguistics, though its executives warn it will take years before machines can truly understand human text.
In the meantime, Powerset may face a bigger challenge: turning its computational break throughs into cash."


Powerset Debuts With Search of Wikipedia - Bits - Technology - New York Times Blog

Powerset Debuts With Search of Wikipedia - Bits - Technology - New York Times Blog: "Ask “Who did Henry VIII marry?” or “What did the FDA ban?” or “What did Bill Clinton sign?” and Powerset will come up with remarkably good answers. (Incidentally, Google does a decent job of answering the first of these questions but not the other two.) Powerset also has other nifty features, like its ability to create mini-dossiers that summarize the information it finds and to take users directly to a section of a document that is most relevant to their search. But Powerset remains a long way off from its promise and faces a seemingly intractable problem: for a very large fraction, if not the vast majority, of searches, keywords work just fine."


Powerset unveils semantic Wikipedia search tool

Article Reuters: "SAN FRANCISCO (Reuters) - Powerset on Sunday unveiled tools for searching Wikipedia that use conversational phrasing instead of keywords, marking the first step of its challenge to established Web search services such as Google.
Powerset's technology breaks down the meaning of words and sentences into related concepts, freeing users from always needing to type the exact words they want to find."