Monday, May 11, 2009

New search engine aims to come up with just the right answer

New search engine aims to come up with just the right answer: "That's the idea behind Wolfram Alpha, a new search service that could be as much of a game-changer as Wikipedia or Google. Alpha, created by renowned mathematician, author, and entrepreneur Stephen Wolfram, uses fast computers and vast statistical databases to answer questions just as a human would — a human with advanced degrees in math."


Thursday, April 02, 2009

Playing with your own MT system...

Various machine translation tools are available like Apertium (GNU license), OpenLogos which is the open source version of Logos Machine Translation System, SYSTRAN which is one of the oldest Machine Translation company and Moses (GNU General Public License).


Moses is a phrase based machine translation tool for converting one language to another language. Technical details regarding this are available on Moses website, www.statmt.org/moses. In this article, is described a short step-by-step process for converting text from one language to another, for example Hindi to English using Moses machine translation tool.

Friday, January 23, 2009

Asia Online Wins Red Herring Global 100 Award for "Most Promising Technology Startups"

"Asia Online's ambition goes far beyond offering world-leading translation technology through its unique online technologies that learn from humans and its crowd-sourcing business model. Its true aim is to make all of the world's knowledge available to every citizen, no matter which language they speak. Called, "the World's largest literacy project," Asia Online is translating tens-of-millions of pages of educational, scientific and historic English-language content into Asian languages.


This content comes from highly valued open sources, such as Wikipedia (the world's seventh most popular online destination), the World Fact Book and tens-of-thousands of published books and essays, and open courseware. Asia Online is also working with publishers of popular English language magazines and publications to translate their materials into Asian languages. This project will effectively eliminate information poverty and provide the same knowledge to every person, regardless of cultural background. On top of this, Asia Online is set to deploy social network services that have proven popular in the West, but have been inaccessible to Asian consumers due to language barriers."


Asia Online's lists as its Chief Scientist Philipp Koehn from the SMT group at the University of Edinburgh.

Saturday, November 29, 2008

Official Google Blog: Our international approach to search

Official Google Blog: Our international approach to search: "... improving Google's international search. This is a tough challenge, since Google search is used in many countries and languages where our engineers have little personal knowledge. Initially, the international search improvements were done by Search Quality engineers who were passionate about their languages and countries: Lina from Sweden improved our parsing of compound words in German and Swedish; Dimitra from Greece introduced diacritical support; Ishai from Israel worked on transliteration corrections for Hebrew and Arabic; Trystan from Australia created methods for identifying local search results and ranking them together with foreign ones from the same language; Alex, a bilingual Ukrainian and Russian, introduced morphological understanding of these languages. As the importance of our international search grew, we solicited help from Googlers in all our offices. Finally, we are leveraging an international network of search specialists who help us understand search within the unique combination of their language and country."

The post is long and there are many improvements presented, however, after running my old tests for Albanian, I see that diacritics still cannot play the discriminating role they should. One of the top pages when searching for 'të' is the page for Tellurium. As suggested in a previous post here 'të' should be entered surrounded by quotes - it's the only way to be assured that the result will contain this exact string. One of the pages returned, containing as many 'ë' as Albanian was in Mixe (Ayuk) - a language spoken in Totontepec, Oaxaca, Mexico.

Tuesday, October 14, 2008

Google goes to the top of the language class - ZDNet.co.uk

Google goes to the top of the language class - ZDNet.co.uk: "Google's machine translation wasn't perfect, but it was well ahead of the competition. On a scale from zero to one, the company's software scored 0.5137 on the Arabic tests and 0.3531 on the Chinese tests. In Arabic, the University of Southern California's Information Sciences Institute came in second with a .4657 and second in Chinese with .3073. IBM scored .4646 on Arabic and .2571 on Chinese."


Sunday, October 05, 2008

Next Week's Turing Test

'Intelligent' computers put to the test | Technology | The Observer: "In the 'Turing test' a machine seeks to fool judges into believing that it could be human. The test is performed by conducting a text-based conversation on any subject. If the computer's responses are indistinguishable from those of a human, it has passed the Turing test and can be said to be 'thinking'.

No machine has yet passed the test devised by Turing, who helped to crack German military codes during the Second World War. But at 9am next Sunday, six computer programs - 'artificial conversational entities' - will answer questions posed by human volunteers at the University of Reading in a bid to become the first recognised 'thinking' machine. If any program succeeds, it is likely to be hailed as the most significant breakthrough in artificial intelligence since the IBM supercomputer Deep Blue beat world chess champion Garry Kasparov in 1997. It could also raise profound questions about whether a computer has the potential to be 'conscious' - and if humans should have the 'right' to switch it off."


Sunday, September 21, 2008

Official Google Blog: The intelligent cloud

Official Google Blog: The intelligent cloud: "We could train our systems to discern not only the characters or place names in a YouTube video or a book, for example, but also to recognize the plot or the symbolism. The potential result would be a kind of conceptual search: 'Find me a story with an exciting chase scene and a happy ending.' As systems are allowed to learn from interactions at an individual level, they can provide results customized to an individual's situational needs: where they are located, what time of day it is, what they are doing. And translation and multi-modal systems will also be feasible, so people speaking one language can seamlessly interact with people and information in other languages."