The Register Home Page

back to article 'Savvy' shortcuts produce near-instant speech-to-speech translation of 36 languages

Meta has developed a machine learning model its researchers claim offers near-instant speech-to-speech translation between around 36 languages. Reminiscent of the Babel Fish from The Hitchhiker’s Guide to the Galaxy, the foundation model SEAMLESSM4T was trained on 4.5 million hours' of recorded human speech and takes a "savvy …

  1. Randy Hudson

    No doubt they’ll use this in WhatsApp and Instagram to better eavesdrop on users dumb enough to give these apps Mic access.

    1. Anonymous Coward
      Anonymous Coward

      Have META cracked AES now?

      1. DS999 Silver badge

        Why would Meta have to have cracked AES to get access to speech data in WhatsApp, an app they wrote?

        1. talk_is_cheap

          Because the encryption takes place at the end device, Meta would have to insert the largest backdoor known to man to access the communications at the end device, which would get noticed rather quickly.

          1. DS999 Silver badge

            They need that large backdoor only if they want to surveil ALL WhatsApp traffic. If they just want to build up some sort of profile of your communications they could reduce the amount by 95% and spit it out across a dozen domains none of which have any obvious connection to Meta, maybe while you are running a browser so you think it is just part of that ordinary back and forth.

          2. Chet Mannly

            Why? They would just do it on the device - the same way that MS does with Skype.

            MS uses Skype to build language models on the local machine, then Skype sends the models to MS for incorporation into the mothership.

            1. Anonymous Coward
              Anonymous Coward

              and it would be discovered so quickly, Musk won't have yet disowned his new best friend of the day.

              1. DS999 Silver badge

                Discovered how?

  2. Ken Hagan Gold badge

    Noisy rooms

    One reason why humans do better than machines in such environments is that humans have two ears but machines typically try to get away with only one microphone.

    Anyone who has ever had a bunged up ear will confirm that noisy environments are much harder to deal with. Presumably the brain uses the second signal to filter out most of the noise, thereby greatly simplifying the problem of understanding. Presumably also, a machine could do the same if it had suitably placed additional microphones. (Most phones already do this, but the placement of their microphones means that it only works when the phone is held like a normal phone by the speaker.)

    1. deive

      Re: Noisy rooms

      You can get binaural microphones for that :-)

      1. Gene Cash Silver badge

        Re: Noisy rooms

        And how many phones have those, again?

        1. deive

          Re: Noisy rooms

          ML isn't just used on mobiles.

          e.g. would be useful for robots, could also be used in in-ear heaphones.

  3. Ken Hagan Gold badge

    Training data

    Unlike some applications, like ChatGPT trained on shitty web pages, this one (speech-to-speech) could presumably acquire a lot of good training data by plugging into the simultaneous translation systems operated by large international organisations like the EU parliament or the UN general assembly.

    I'm sure those orgs would be happy to co-operate in exchange for a suitable payment, since they might also be among the main beneficiaries down the line.

    1. Roland6 Silver badge

      Re: Training data

      Also that must have been some Audible bill !

  4. An_Old_Dog Silver badge

    Hahhhh-Hahhhh

    1. Anyone with a knowledge of one or more foreign languages and who have watched media with English subtitles while hearing dialogue in one of the foreign languages they know can tell you the translation quality varies from good to hilariously bad.

    2. Further muddying the waters are the companies using LLMs to generate (low-quality) subtitles.

    Subtitles are thus poor-to-bad-quality LLM training material.

    1. Ace2 Silver badge

      Re: Hahhhh-Hahhhh

      “Ne joue pas le fou!”

      Subtitle: Don’t fuck up!

    2. Anonymous Coward
      Anonymous Coward

      Re: Hahhhh-Hahhhh

      While I agree in general, remember how much people laughed at google translate when it appeared several years ago, and businesses around the world jump at this opportunity, with hilarious effect. Fast forward several years later, those google, yandex, deepl translations are far from bad. They still fail miserably against spoken language and - for now - it takes less time to translate from scratch, using human translator, than to 'repair' the same audio file translated by machine. The balance of time and cost favours humans for now. As it did with written texts before automatic text translation engines got past and above 'good enough'. So, most translation jobs now are post-machine-translation-edition, "because cheap(er)". Speech is much harder to crack, but 'cheaper' is a driving force for 'progress'.

    3. JLV Silver badge

      Re: Hahhhh-Hahhhh

      Yeah, I see what you mean:

      Les Inconnus, Je t'embete https://www.youtube.com/watch?v=VtBlELjpHAM

      (damn! 240p, but looks more like 64p in quality)

  5. MOH

    Where did the 4.5 million hours of spoken audio come from?

    Facebook parent you say?

  6. Ol'Peculier

    Mostly Harmless?

    "Meanwhile, the poor Babel fish, by effectively removing all barriers to communication between different races and cultures, has caused more and bloodier wars than anything else in the history of creation"

    What with that and Trumpton, interesting times ahead...?

  7. Anonymous Coward
    Anonymous Coward

    There is hope

    One day we may be able to finally understand the Welsh.

    1. Primus Secundus Tertius

      Re: There is hope

      The Romans just insisted that the barbarians should use Latin.

  8. Primus Secundus Tertius

    No worse than humans

    "The tool also struggles in many situations that humans handle with relative ease"

    You are exaggerating human abilities here. Very likely the tool is no worse than the average (as opposed to the best) human.

    1. Chet Mannly

      Re: No worse than humans

      You are exaggerating the tool's abilities. Every machine translator I've used makes mistakes and mistranslations a native 6th grader would get right.

  9. Anonymous Coward
    Anonymous Coward

    A few loose thoughts

    "Importantly, we believe that SEAMLESSM4T-fueled applications should best be viewed as an augmentation device that assists in translation rather than a tool that replaces the need for language learning or reliable human interpreters."

    - they're not stupid so I consider them disingenuous: it will be used EXACTLY as a "tool that replaces the need for language learning or reliable human interpreters". (btw, I'm not an interpreter).

    '...although we believe that language acquisition should remain a key mechanism for boosting our world-readiness, we acknowledge that doing so requires resources many people may not possess'.

    - it's a lie, most humans on this planet 'may' and DO possess, practically FREE, resources to acquire a foreign language. I mean, sure, you need access to electricity, laptop and internet, but even access to those can be found for free, and not only in the West.

    "Starting with some data that they knew to be reliable..."

    - how do they 'know' what is 'reliable'? Never mind the subjectivity of 'reliable', when 2 decent translators can't agree which version is correct, and which one is only 'ok-ish', but where do you take 'reliable' pairs from? "Video clip and corresponding subtitle" ok... so surely not youtube, because you can't distinguish which video+subtitles pair is 'reliable', which one is 'human crap' and which one is 'AI crap' and which one 'passable', unless, knowing both languages, you analyse the pair, and still it's going to be a subjective decision. The only dataset I can think of, which gives SOME degree of reliability would be mainstream movies (never mind occasional fails), or, perhaps, I don't know, UN interpretation files, if that set is available for research, that is. They mention the UN corpus for written translations only, and I can't see in the abstract what exactly their 'reliable source' for speech was, only this:

    'we created a corpus of more than 470,000 h of automatically aligned speech translations'.

    - also, I'm kind of sceptical about 'near-instant' claim (again, define 'near-instant'). If a key element appears at the end of a sentence / statement you need to wait for the sentence to finish to find the closest match, even though (I imagine) the processing and narrowing is being done on the fly.

    Also:

    "...cultural reasons (that is, to become a global citizen)..."

    - I think they've missed recent trends on the subject of 'global village' and 'global citizens', and by 'recent' I don't specifically mean Trump. That said, globalisation is unlikely to be stopped even with growing fragmentation. And then, we can interrogate 'foreign combatants' much more effectively, here's a cheque!

    Overall, progress is gooood (those fpv-rpg drones are cooler than machine gun, eh?) though ultimately, another step in the process of stupefying humans.

    1. Chet Mannly

      Re: A few loose thoughts

      "how do they 'know' what is 'reliable'?"

      100% this. Skype builds language models locally for incorporation into MS' larger models. These local models are built on my language lessons in very rudimentary Italian and somewhat better Spanish. I can't imagine how horrific the models Skype is building based on my terrible 2nd and 3rd languages are...

  10. LVPC

    >> .although we believe that language acquisition should remain a key mechanism for boosting our world-readiness, we acknowledge that doing so requires resources many people may not possess'.

    More than 40% of the world speaks 2 or more languages. And it has zero to do with "world readiness." Most countries have significant groups that use a second, or even 3rd, language.

    These researchers need to get into the real world more, and do some real-world research.

    1. Chet Mannly

      They do lots of real-word research - in California :)

POST COMMENT House rules

Not a member of The Register? Create a new account here.

  • Enter your comment

  • Add an icon

Anonymous cowards cannot choose their icon

Other stories you might like