sexta-feira, 15 de junho de 2012
Babel or babble
Mr Ostler renewed his claims that machine translation does away with the need for a lingua franca at the Hay Literary Festival last week. Many linguists disagree with him, including David Crystal, who has forecast that English may “find itself in the service of the world community forever.” But we once thought the same about Greek, Persian, Latin and French, all of which became obsolete.
“Evidently, automatic systems replacing a real lingua franca is likely to be a bit different” than one lingua franca replacing another, Mr Ostler told your correspondent. “The process will run faster for mainstream languages than for the peripheral smaller fry. But I think the development will be unmistakable within one generation—say by 2050.”
Yet the quality of machine translation is still uneven at best. Google Translate is probably the best of the bunch. It has 200m users a month by the last count, and translates in a single day the equivalent of the whole annual output of the globe’s professional translator corps. While there’s no doubt that it is a powerful and useful tool, it can produce embarrassing errors. There are fewer mistakes than there were in 2001 when Google launched it, but there still seems a long way to go until sci-fi technology becomes real.
It’s feasible that machine translation could replace human translators for written texts. Text is easier to translate than conversation, and is better suited to the technology, which is "trained" by huge corpora of human-translated texts. But these tools are only as good as the corpora themselves. Google Translate draws on a variety of texts, including documents from the European Union and United Nations. These are chock-full of legalese, and rarely represent the everyday language of man. Finding more representative texts without raising the hackles of copyright enforcers could prove a stumbling block.
There is also a remaining reliance on English as an intermediary language for translation. Precious few Galician novels will have been translated into Welsh nor Welsh novels into Galician, but both make their way into English. Though humans may not use English as a lingua franca in the future, necessity dictates that the machine translators still will (as Asya Pereltsvaig seems to have busted Google Translate doing). What machine translation will allow, says Mr Ostler, is those “minority language speakers to intervene—at last—directly in wider issues to a wider audience” as their niche languages become easily translatable by gadgets. That does, of course, discourage people from learning the language in the first place.
Spoken language is too quick and fragmented for machine translation today. Simple dictation software must be carefully trained, never mind the extra step of translation. Ordinary conversation is full of false-starts and errors that speakers and listeners barely notice, but which computers are baffled by. And of course tone of voice, cultural references, idiom and humour multiply the challenges. The true-to-life Babel fish is some ways away.
Idealist and pessimist tribes gather around the future of machine translation. Your correspondent is a pessimist, but would dearly like to be proved wrong, because Mr Ostler's vision of the future is an exciting one. “Increasingly, electronic tools will be there for people to get what they need from documents and recordings in foreign languages,” he explains. “And real-time gizmos will support people’s needs in face-to-face interaction too.” Verb conjugation tables and vocabulary lists may be consigned to history. Add schoolchildren to the many others who would rejoice in a lingua-franca-free future.
The Economist
segunda-feira, 28 de fevereiro de 2011
8th International Workshop on Natural Language Processing and Cognitive Science - NLPCS 2011
http://www.cbs.dk/nlpcs2011
20-21 August, 2011 - Copenhagen, Denmark
Important Dates for NLPCS 2011
Paper Submission: 2nd May 2011
Authors' Notification: 6th June 2011
Final Paper Submission: 15th July 2011
Registration and Payment: 15th July 2011
Workshop : 20-21 August 2011
Scope and Topics of NLPCS workshop
The aim of this workshop is to foster interactions among researchers and practitioners in Natural Language Processing (NLP) working within the paradigm of Cognitive Science (CS). Research into NLP involves concepts and methods from many fields including artificial intelligence, linguistics, computational linguistics, statistics, computer science, and most importantly cognitive science.
The overall emphasis of the workshop is on the contribution of cognitive science to language processing, including conceptualisation, representation, discourse processing, meaning construction, ontology building, and text mining.
The special theme of this year's NLPCS workshop is "Human-Machine Interaction in Translation" . Therefore, we particularly welcome papers addressing aspects of human and machine translation and human-computer interaction in translation.
Additional topics of interest include, but are not limited to:
– Cognitive and Psychological Models of NLP
– Computational Models of NLP
– Evolutionary NLP
– Situated (embodied) NLP
– Multimodality in speech / text processing
– Text Summarisation and Information Extraction
– Natural Language Interfaces and Dialogue Systems
– Multi-Lingual Processing
– Pragmatics and NLP
– Speech Processing
– Tools and Resources in NLP
– Human and Machine Translation
– Ontologies
– Text Mining
– Electronic Dictionaries
– Evaluation of NLP Systems
These topics can be addressed from any of the following perspectives:
full automation by machines for machine (traditional NLP or HLT), semi-automated processing, i.e. machine-mediated processing (programs assisting people in their tasks), simulation of human cognitive processes.
With this year's special theme we also welcome submissions on translators' experiences with CAT tools, human-machine interface design, methods for and evaluation of interactive machine translation, feasibility studies, user simulation, etc.
segunda-feira, 24 de janeiro de 2011
Search giant admits translations can be lacking
Google vice-president Vint Cerf said the use of statistical translation methods, where translations are made on the basis of probabilities and don't rely on parsing, had vastly improved online translation.
But he warned about their reliability and said there were problems with interpreting the meaning of the same phrase in British and American English, let alone phrases in different languages.
"I'd be really careful about having any kind of a sensitive debate with someone either spoken or written using these translations," he said during a visit to The Australian's Sydney bureau last week.
Google offers online translations between more than 50 languages. It employs voice recognition in the Google iPhone app. It is seeking to combine these technologies to produce a phone capable of real-time voice translation.
Dr Cerf said Google has begun building translation statistics by comparing phrases found in documents produced in multiple languages.
"Now our research teams are starting to combine the statistical methods, which work well when we have a large body of translated materials to work with, to multiple documents written in different languages that purport to be the same thing.
"Our national institutes of standards and technology have rated the Google translations as above all others that I know about.
"If we were going from zero to 10, we would be about five, that's better than almost everybody else.
"But I can tell you that I read newspapers from other countries by using Google Translate and at least I'm getting a pretty good gist of what's being said and if I need to know more I'd go to a language speaker, an expert speaker."
The Australian
quarta-feira, 13 de janeiro de 2010
Um tradutor quase humano
O ultramoderno laboratório da IBM em Yorktown Heights, bucólico subúrbio do estado de Nova York, foi projetado nos anos 50 pelo arquiteto finlandês Eero Saarinen. Desse complexo saem algumas das ideias tecnológicas mais ousadas do mundo. Uma delas promete mudar a face da colaboração entre equipes multinacionais e multiétnicas nos negócios da IBM. O seu nome é n.Fluent (pronuncia-se én-flúãnt). Desde que foi anunciado, em novembro último, o n.Fluent coleciona elogios. Trata-se de um software de tradução em tempo real do inglês para outras dez línguas. Com ele, a IBM será capaz de verter de forma acurada o conteúdo de páginas da web, documentos eletrônicos, chats e mensagens instantâneas. Ainda em teste, o n.Fluent funcionará também em smartphones.
A questão da tradução linguística para incentivar a colaboração foi eleita uma das dez prioridades da empresa. “Para a inteligência planetária, o mundo precisa de um vocabulário comum à colaboração, em especial para a comunidade de negócios”, diz David Lubensky, pesquisador da empresa e um dos gestores do projeto. O n.Fluent pode ser uma mão na roda. A IBM possui 400 mil funcionários em 170 países e o compartilhamento crescente de tarefas e projetos entre esses países cria o risco de uma Babel. O n.Fluent começa a ser apontado como o antídoto para o problema. O software estará disponível para a tradução do inglês ao chinês (simplificado e tradicional) e ao coreano, japonês, francês, russo, alemão, espanhol, português, italiano e árabe. O inglês é o parâmetro e não existe, pelo menos até agora, módulos de tradução entre as outras dez línguas – do português para o russo, por exemplo.
Leia mais na Época Negócios
terça-feira, 24 de março de 2009
Texto exibido na ferramenta de tradução automática da Microsoft
http://www.microsofttranslator.com/
domingo, 22 de março de 2009
Tradutor online da Microsoft está disponível no Windows Live Messenger e na SciELO
Desenvolvida pela Microsoft, a ferramenta de tradução online também está presente em diversos aplicativos da companhia, como o Internet Explorer 8, Word 2007 e o Live Search. "Com o Windows Live Translator, milhares de usuários têm a oportunidade de acessar páginas da web em inglês ou em outras línguas, além da conveniência de iniciar uma conversa e contar com ajuda para traduzi-la em diferentes idiomas", afirma Carolina Aranha, Gerente geral da divisão Online da Microsoft Brasil.
http://www.microsoft.com/latam/presspass/brasil/2009/marco/bot.mspx