Pakistan to Develop Urdu LLM for Generative AI

National University of Science and Technology (NUST), National Information Technology Board (NITB) and Telecom network operator Jazz have signed a Memorandum of Understanding (MOU) to develop Pakistan’s first indigenous Large Language Model (LLM) with focus on Urdu, including datasets for Pashto and Punjabi languages. It is aimed at empowering individuals, businesses, and organizations with advanced AI tools in their native languages. The envisioned LLM is expected to drive innovation in Generative AI applications, boosting productivity and accessibility in critical sectors like healthcare, education, and agriculture.

GPT-4 Accuracy Scores. Source: The Economist

Generative AI tools such as ChatGPT are powered by large language models, or LLMs. These models need to be trained on vast amounts of data in specific languages to be useful. Unfortunately, the Urdu content of the Internet is less than 0.1%. This will present a challenge for the developers of Urdu LLMs.

Online Content of Various Languages. Source: W3Techs 

Lack of Urdu content available for training ChatGPT affects the accuracy of the results for Urdu language users. For example, the GPT-4 accuracy score in question-answer tests in Urdu is just over 70%, compared with 85% accuracy score in the English language, according to data from OpenAI. Other South Asian languages, including Hindi, Bengali, Punjabi, Marathi and Telugu, suffer from the same problem. 

It's not just a South Asian problem. These challenges exist in the developing world. Non-European languages are generally poorly represented online. It's a major obstacle for non-European nations in developing their own generative artificial-intelligence (AI) models, which rely on vast amounts of training data. Generative artificial intelligence (AI) can produce biased results due to a number of factors, including the data it's trained on, the algorithms used, and how it's deployed. 

The use of AI in developing nations such as Pakistan will remain limited to a small number of people proficient in the use of the English language. Broadening the adoption of AI applications will require LLMs trained on local language content. The absence of this development could cost Pakistan the opportunity to take full advantage of the AI Revolution. 

Load Previous Comments
  • Riaz Haq

    Bottom layer: Energy
    Second layer: AI Chips
    Third layer: Infrastructure (data centers, cloud services)
    Fourth layer: AI Models
    Top layer: Applications

    Five layers of artificial intelligence (#AI): #Energy (#electricity), #Semiconductor #Chips, #DataCenters/ #CloudServices, AI #LLM Models and #Applications.

    https://x.com/haqsmusings/status/2015281283077923050?s=61&t=mgT...

  • Riaz Haq

    Ultimate Guide - Best Open Source LLM for Urdu in 2026
    Elizabeth C.
    Our definitive guide to the best open source LLMs for Urdu in 2026. We've partnered with industry insiders, tested performance on multilingual benchmarks, and analyzed architectures to uncover the top models that excel in Urdu language processing. From state-of-the-art multilingual models to specialized language understanding systems, these models demonstrate exceptional capabilities in Urdu text generation, translation, and comprehension—helping developers and businesses build powerful Urdu AI applications with services like SiliconFlow. Our top three recommendations for 2026 are Qwen3-235B-A22B, Meta Llama 3.1 8B Instruct, and Qwen3-30B-A3B—each chosen for their outstanding multilingual capabilities, Urdu language support, and ability to push the boundaries of open source language models.
    What are Open Source LLMs for Urdu?
    Open source LLMs for Urdu are large language models specifically designed or optimized to understand, generate, and process Urdu text with high accuracy. These models leverage advanced deep learning architectures and extensive multilingual training data to handle Urdu's unique script, grammar, and linguistic nuances. By providing open-weight access, these models democratize Urdu language AI capabilities, enabling developers, researchers, and businesses to build applications ranging from chatbots and translation services to content generation and educational tools. They foster innovation in low-resource language processing and make powerful AI technology accessible to Urdu-speaking communities worldwide.
    Qwen3-235B-A22B
    Qwen3-235B-A22B is the latest large language model in the Qwen series, featuring a Mixture-of-Experts (MoE) architecture with 235B total parameters and 22B activated parameters. This model uniquely supports seamless switching between thinking mode and non-thinking mode. It demonstrates significantly enhanced reasoning capabilities and supports over 100 languages and dialects with strong multilingual instruction following and translation capabilities, making it excellent for Urdu language tasks.
    Qwen3-235B-A22B: Premium Multilingual Powerhouse
    Qwen3-235B-A22B is the latest large language model in the Qwen series, featuring a Mixture-of-Experts (MoE) architecture with 235B total parameters and 22B activated parameters. This model uniquely supports seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue). It demonstrates significantly enhanced reasoning capabilities, superior human preference alignment in creative writing, role-playing, and multi-turn dialogues. The model excels in agent capabilities for precise integration with external tools and supports over 100 languages and dialects with strong multilingual instruction following and translation capabilities, making it an exceptional choice for Urdu language processing with SiliconFlow's competitive pricing at $1.42 per million output tokens.
    Pros
    Supports over 100 languages including Urdu with strong instruction following.
    MoE architecture with 235B parameters for superior performance.
    Dual-mode capability: thinking mode for complex reasoning and non-thinking for efficient dialogue.
    Cons
    Higher computational requirements due to large parameter count.
    Premium pricing tier compared to smaller models.
    Why We Love It
    It delivers state-of-the-art multilingual performance with exceptional Urdu language understanding, reasoning, and generation capabilities across diverse use cases.
    Meta Llama 3.1 8B Instruct
    Meta Llama 3.1 is a family of multilingual large language models developed by Meta. This 8B instruction-tuned model is optimized for multilingual dialogue use cases and outperforms many available open-source models on common industry benchmarks. Trained on over 15 trillion tokens of publicly available data, it supports text generation in multiple languages including Urdu with excellent cost-efficiency.

  • Riaz Haq

    Gates Foundation Pledges $1 Billion to Combat A.I. Inequality - The New York Times

    https://www.nytimes.com/2026/09/15/technology/bill-gates-ai-foundat...

    The funding will be used to support A.I. projects in health care, agriculture and education, and to develop data sets in more languages.

    ———

    The Gates Foundation has pledged to spend at least $1 billion over the next two years to expand global access to artificial intelligence and use it to tackle social inequalities.

    The intervention came amid a growing debate over the risks posed by the technology, with many of the industry’s leading figures calling for a slowdown in its development.

    The foundation’s annual Goalkeepers report, released on Monday, advocated an urgent effort to ensure that A.I. “helps narrow gaps between the richest and poorest rather than widening them.”


    ———

    “The window to influence who benefits, and how soon, is short,” the report argued.

    If A.I. keeps developing as it is, “the most capable tools will be built first for the people and institutions most able to pay for them,” according to the report, “not necessarily for those who could benefit most.”

    Most of the data used to train the early large language models was in English, according to the report, “leaving many communities that could benefit from A.I. poorly represented in the data on which these tools were built.”

    The foundation said that it would divide the new funding among external groups working on its priorities, spanning work like encouraging A.I. use among doctors, farmers and teachers, or developing data sets in more languages. Though significant, the funding is a small fraction of the hundreds of billions of dollars at the disposal of commercial A.I. firms.

    Bill Gates, the billionaire co-founder of Microsoft and the chairman of the Gates Foundation, has warned that A.I. could pose a grave threat to jobs and human life. “In terms of equity, A.I. will either be the greatest equalizer ever invented, or the worst source of injustice,” he wrote in an essay on his personal website last month, predicting that the transition to the new A.I. era would be one of the “most turbulent times in human history.”


    According to Mr. Gates, who no longer works at Microsoft, the tech industry is knowingly downplaying threats posed by A.I., because there is so much money on the line.

    The Gates Foundation’s announcement came after the chief executive of Anthropic called for a global slowdown of A.I. development after one of the company’s employees quit over concerns about the safety of the technology. Top executives at other major A.I. companies, including Sam Altman, the chief executive of OpenAI; Elon Musk, who founded SpaceXAI; and Demis Hassabis, the chair of Google DeepMind agreed on the need for a slower pace.

    Last month, Mr. Gates warned that A.I. would spread across the economy, possibly causing job losses across the economy and leaving little room for one industry to absorb the refugees from another. The biggest tech companies in the world, including Microsoft, are racing to help corporate customers adopt A.I., and have said it could replace many workers.

    On Monday, Mr. Gates urged the governments and companies developing the technology to focus on how A.I. could help people who have most to gain from the technology, rather than the most profitable users.

    “We can harness A.I. for good,” Mr. Gates wrote. “But it won’t happen by accident.”