Indonesian researchers and AI experts have been on a mission to document and preserve native languages and local dialects across the archipelago. However, with sparse funding and support, they have opted for creating a model that can stand as the blueprint for future projects.
In just three years, the team behind a dialogue summarization dataset collected more than 10,000 dialogue and summary pairs, with a total of 3 million words for three underrepresented languages: Balinese, Buginese, and Minangkabau.
The project, dubbed NusaDialogue, covers 17 topics and 185 subtopics, such as history, religion, hobbies, culture, electronics, and food, along with annotations provided by 73 native speakers across the three regions.
Co-designed by Indonesian AI company Prosa.AI in collaboration with the FAIR Forward initiative by the Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ), the dataset is publicly available for anyone to cite and download through open-source platform Hugging Face and FAIR Forward’s Catalog of Open Source AI Datasets, Models & AI Systems.
However, not many people outside of the sphere of researchers and developers are aware this project exists.
This is largely due to the fact that the project does not reach the end user. The team made a conscious decision not to implement it into a chatbot or a ready-for-use commercial product.
“Few are willing to invest in the research. Not many people realize that creating AI involves a long journey of data collection, especially when it is done ethically”
Karlina Octaviany

Digital Anthropologist and Artificial Intelligence Advisor at (GIZ) GmbH, Karlina Octaviany, who was directly involved in the NusaDialogue project, explained that oftentimes people only want to see the “flashy things in the end,” but they fail to grasp the steps necessary to build such projects.
“People don’t quite understand the ‘garbage in, garbage out’ principle when it comes to AI. If the dataset is poor, the AI platform’s output will be poor, too. That’s precisely why we want high-quality data for AI,” stated Octaviany in an interview with The Sociable.
As for whether there are plans to expand the data into a chatbot, Octaviany stated plainly that there aren’t any. This is because the funding granted for the program was only enough to build a comprehensive dataset and not a server that can host the implementation of the data.
Octaviany likens the project to a blueprint for a “foundation that can be used to build houses.”
“The most expensive part is actually acquiring datasets for research, to train the models. Yet, few are willing to invest in the research. Not many people realize that creating AI involves a long journey of data collection, especially when it is done ethically.
“It is precisely this research and development process [of AI] that rarely secures funding,” said Octaviany.
According to Octaviany, FAIR Forward and Prosa.AI made a project that would be sustainable in the long run, in hopes that many other Language Learning Models can spawn from it.
So far, the little-known language model has been downloaded 56 times on Hugging Face. Meanwhile, the Nusa Translation dataset has been downloaded only 20 times.
The FAIR Forward program was only meant to go on for three years, funding not only the NusaDialogue project, but also a range of others meant to democratize the use of Artificial Intelligence in Indonesia.
Another project that was implemented alongside Prosa.AI included FaktaIklim (Climate Facts), an open-source, AI-powered platform and chatbot in Indonesia designed to detect and counter climate change misinformation.
Unlike NusaDialogue, FaktaIklim is sponsored by the government, mainly through the Indonesian Agency for Meteorology, Climatology, and Geophysics.
It was during the closing ceremony for FAIR Forward’s initiatives in May 2026 that Octaviany found out from Prosa.AI’s then-Natural Language Processing Lead Ayu Purwarianti that Prosa.AI had been acquired by a larger company and would be closing down.
As saddened as she was that their long-time development partner would cease to exist, Octaviany was relieved that the project had been built to last, with or without Prosa.AI’s continued operation.
“It’s not managed by Prosa anymore. The dataset is now public. So, it’s not on Prosa’s servers either. All the datasets have already been uploaded to GitHub for public use,” she said.
The fate of Prosa.AI after NusaDialogue
Lecturer at Bandung Institute of Technology and ex-Natural Language Processing (NLP) Lead from Prosa.AI, Ayu Purwarianti, confirmed that the AI company had been acquired by GLAIR (GDP Labs AI Research), a subsidiary of Indonesian conglomerate Djarum Group.
“Yes, yes it was [acquired]. All the resources that Prosa owned now belong to GLAIR,” she stated.
Some Prosa.AI employees stayed after the acquisition, while others found other opportunities in the AI and development sector.
“Most of the Prosa alumni are doing relatively well in the current job market. Because their skill set is still rare. So, it’s not hard for them to find a new job.”
“Startups aren’t given special treatment; they are treated just like any other company. So, when a startup lacks sufficient capital, it’s a real struggle”
Ayu Purwarianti

Prosa.AI was one of the first home-grown AI startups in Indonesia. It was established in 2018, years before the AI boom hit global markets. They have been the brains behind various products based on artificial intelligence and natural language processing tailored for the Indonesian language.
Some notable products by Prosa.AI included Prosa Text-to-Speech (TTS), Prosa Meemo for automated meeting transcription, and Prosa Conversa for customer service chatbots.
Reflecting on her experience with Prosa.AI, Purwarianti explained how it is still very difficult for local startups in Indonesia to compete with big tech companies from abroad, especially giants like OpenAI and DeepMind.
Most of the time, big foreign AI companies offer their services for free while local Indonesian companies rely on subscription-based models to stay afloat.
“You have foreign players like Google and the likes. Since they are major corporations, they offer more comprehensive packages, going beyond just speech-to-text by including a wide range of additional features.
“We simply can’t compete [with foreign companies]; our packages don’t measure up. Domestically, there isn’t much enthusiasm for adopting locally made solutions, unfortunately,” she said.
The government also has not given any incentives to bolster the growth of local AI startups. There are currently no tax exemptions or funding initiatives that are enough to keep servers running.
“Startups aren’t given special treatment; they are treated just like any other company. So, when a startup lacks sufficient capital, it’s a real struggle,” said Purwarianti.
At the moment, the Indonesian Ministry of Communication and Digital Affairs is planning to release two regulations to boost the AI startup industry, which are the Artificial Intelligence National Roadmap and the Guidelines for AI Ethics.
Director General of Digital Ecosystems at the Indonesian Ministry of Communication and Digital Affairs, Edwin Hidayat Abdullah, stated that digital regulations must be able to support the growth of local digital startups.
According to Abdullah, the government should position regulations as catalysts that foster innovation, provide certainty for investors, and accelerate the development of startups as well as Indonesia’s digital economy.
“Regulations should not only drive compliance; they must also act as catalysts for digital economy growth,” said Director General Edwin in a July 2022 press release.
Although Prosa.AI is no longer in business, Purwarianti hopes that NusaDialogue will continue to inspire more developers in the Indonesian AI ecosystem to work on AI initiatives that preserve native languages.
“Our goal was that this could serve as a trigger point. To encourage other researchers to take action and start gathering data in their respective regions to build language datasets. But, as it turns out, we have not quite accomplished that,” said Purwarianti.
Image Source: FAIR Forward, courtesy of Karlina Octaviany
