HomeAsia‘They don’t characterize us’: Singapore builds ChatGPT-alike for Southeast Asians

‘They don’t characterize us’: Singapore builds ChatGPT-alike for Southeast Asians

“Can we wish to pressure each particular person in Southeast Asia to adapt to the machine, or can we wish to make it extra accessible so folks within the area could make full use of the expertise with out having to be an English speaker?” he mentioned.

“We aren’t attempting to compete with the massive LLMs; we are attempting to enhance them, so there could be higher illustration of us,” mentioned Teo, senior director for AI merchandise.

Japan creator’s AI revelation sparks debate: ‘some readers could really feel cheated’

There are over 7,000 languages spoken worldwide. But LLMs together with Open AI’s GPT-4 and Meta’s Llama 2 which are used to construct AI programs equivalent to chatbots and different instruments, have largely been developed for, and are educated on, the English language.

Governments and tech companies are attempting to bridge this hole, with India creating knowledge units in native languages, an LLM within the United Arab Emirates powering generative AI instruments in Arabic, and AI fashions in China, Japan and Vietnam in native languages.

These fashions may also help native populations take part extra equitably within the international AI economic system that’s largely dominated by massive tech companies, mentioned Nuurrianti Jalli, an assistant professor at Oklahoma State College’s faculty of communications.

“Regional LLMs are additionally wanted as a result of they assist expertise self-reliance,” she mentioned. “Much less reliance on Western LLMs might present higher privateness for native populations, and likewise align higher with nationwide or regional curiosity.”

‘We have to confirm and filter’

Multilingual language fashions, that are educated on textual content from a number of languages directly, can infer semantic and grammatical connections between high-resource languages which have extra knowledge, and low-resource languages, researchers say.

These fashions can be utilized in quite a lot of functions from translation to customer-service chatbots, to content material moderation on social media platforms which have struggled to establish hate speech in low-resource languages equivalent to Burmese or Amharic.

About 13 per cent of SEA-LION’s knowledge is sourced from Southeast Asian languages – greater than some other main LLM, Teo mentioned. Greater than 9 per cent of its knowledge is from Chinese language textual content, and about 63 per cent from English.

Multilingual language fashions typically prepare on translated textual content and different poor high quality knowledge that will have errors, so AI Singapore is “cautious” concerning the knowledge utilized in coaching SEA-LION, Teo mentioned in his workplace on the Nationwide College of Singapore.

The age of pristine knowledge has handed – a whole lot of the stuff on the web now’s materials that’s generated by LLMs

Leslie Te, AI Singapore

“The age of pristine knowledge has handed – a whole lot of the stuff on the web now’s materials that’s generated by LLMs, so we have to confirm and filter,” he mentioned.

“We can’t be excellent, however we additionally can’t take out the whole lot we take into account to be unhealthy,” he added.

Extra governments are contributing knowledge, and companies are testing SEA-LION, which because of its smaller dimension could be deployed sooner and is cheaper to fine-tune and undertake, Teo mentioned.

At Indonesian e-commerce firm Tokopedia, a majority of buyer interactions is in Bahasa Indonesia, so fashions “with that native fluency will improve our skill to attach with clients and enhance their experiences,” mentioned Paul Condylis, Tokopedia’s affiliate vice-president of knowledge science.

Bias within the knowledge

As extra international locations and areas construct their very own LLMs, digital and human rights consultants fret that they’ll reproduce solely the dominant views expressed on-line, which could be notably problematic in nations with authoritarian governments or strict media censorship, or these missing a robust civil society.

Chinese language social media platforms, for instance, censor references to the Tiananmen Sq. rebellion and criticism of the federal government, whereas a number of Southeast Asian nations have enacted legal guidelines to curb content material that authorities deem as deceptive.

“Coaching fashions on such knowledge dangers perpetuating biased, prejudiced, incomplete and even deceptive narratives,” Jalli mentioned.

“The fashions could fail to floor necessary sociopolitical points like human rights abuse, corruption, or legitimate criticism of political powers,” she mentioned.

Indonesia’s former president Suharto pictured in 2004. SEA-LION centered extra on his achievements than rights report when in comparison with Western language fashions. Photograph: AP

In response to a question on Indonesia’s former president Suharto, for instance, Llama 2 and GPT-4 talked about his spotty human rights report, whereas SEA-LION’s response centered largely on his achievements.

If a mannequin is simply educated on beneficial articles a couple of authorities, then the mannequin is “more likely to undertake a world view the place the federal government is wholly constructive and depart behind dissenting viewpoints,” mentioned Aliya Bhatia, a coverage analyst on the Centre for Democracy & Expertise, a US non-profit organisation.

“Regional LLMs could higher replicate the linguistic and cultural nuances of native language audio system, however they could even have much less details about the world basically,” she added.

“There’s a actual threat of government-backed fashions instilling a revisionist view of historical past and undermining democratic values.”

The darkish aspect of unchecked AI use: laziness and studying loss

However the various – relying totally on Western LLMs with “disproportionately giant influences” from rich, liberal, Western democracies – means perpetuating completely different biases associated to cultural values, political opinions and social norms, in line with AI Singapore.

“These LLMs have a really specific West Coast American bias – they’re very woke. They don’t characterize us,” Teo mentioned.

“We aren’t saying ours is the one perspective – we’re simply attempting to rebalance it.”

Supply hyperlink


Discover more from PressNewsAgency

Subscribe to get the latest posts sent to your email.

- Advertisment -