Showing posts with label MachineLearning. Show all posts
Showing posts with label MachineLearning. Show all posts

Epic history of LLM


RNN. Seq to seq NLP tasks. 

1. Many to one: Sentimental Analysis

2. One to Many: Image caption

3. Many to Many: 

- Synch many to many: # input = # output. E.g. Part of speech tagging, Named Entity Recognition

- Asynch many to many: translation, text summarization, question and answer, chatboat, speech to text, 

Seq2seq model is used for Many to Many

Stage 1: 2014 Encoder decoder network

Encoder and decoder are LSTM. RNN and GRU are other options. 

It is good for small sentences. Not for 30+ words

BLEU score

Stage 2: 2015 Attention Mechanism

Encoder is same

Attention Mechanism: Attention layer at decoder finds out which hidden state is useful at each stage of decoder and generate context vector for that stage. So, Multiple context vectors based on  encoder's (hidden state of LSTM = ctht vector) are available to decoder. 

Training time is more. 

2015 to 2017: May types of Attention Mechanisms were introduced. 

Stage 3: 2017 Transformer

No LSTM

No RNN Cell

Self-attention was introduced

Both encoder and decoder uses attention

Transformer can process all words in parallel 

1. Attention layer = Multi Head Attention

2. Normalization Layer

3. Dense Layer

4. Input embeddings

It needs hardware, time, and data

Stage 4: 2018 Jan Transfer Learning

Challenges

1 Single model cannot perform all tasks like sentimental, translation, summarization

2 lots of labeled data

Universal Language Model Fine-tuning ULMFiT proposed to use Language modelling  as Pre-training. Language modelling is NLP task to predict next word. Advantages

1. Rich feature training

2. unsupervised task

model: AWD LSTM model

data set: wikipedia

finetuning changed output as classifier with many data set 

Scratch 10000 data. Now fine tune 100 data still better result

- No transformer

Now in 2018, we have two technolgoies

1. architecture: transformer

2. training. Pretrain and transfer learning

Stage 5: 2018 Oct LLM

Transfer learning on transformer

1. Google : BERT (encoder only model) 

2. OpenAI: GPT (decoder only model)

LM to LLM

1. data

2 hardware GPU clusters

3 time : days to weeks

4. cost =  h/w + electricity + people + infra

5. energy consumption 

---------------

GPT3 - > chatGPT

1. RLHF : Reinforcement Learning from Human Feedback

2. incorporate safety and ethical guidelines 

3. improvement in contextual point

4. dialogue specific 

5. continuous improvement based on user feedback


Reference https://www.youtube.com/watch?v=8fX3rOjTloc&list=PPSV

DSPy


DSPy = Declarative Self-improving Python.

Components

1. language model — LLM that will answer our questions,

2. signature —a declaration of the program’s input and output (what task we want to solve),

 - 1. inline

 - 2. class

dspy.InputField()

List[Literal['', '', '']] = dspy.OutputField()

3. module — the prompting technique (how we want to solve the task).

 - Building blocks

 - different prompting strategies, 

  - 1. dspy.Predict

  - 2. dspy.ChainOfThought 

  - 3. dspy.ReAct (to add tools = function calling

4. Optimiser

  - 1. Automatic few-shot learning (e.g. BootstrapFewShot or BootstrapFewShotWithRandomSearch)

  - 2. Automatic instructions optimisation (e.g. MIPROv2) 

  - 3. Automatic fine-tuning (e.g, BootstrapFinetune) 

Other points

  • dspy.inspect_history for logs
  • Caching

# 1. updating config

dspy.configure_cache(enable_memory_cache=False, enable_disk_cache=False)

# 2. not using cache for specific module

math = dspy.Predict("question -> answer: float", cache = False)

  • dspy.configure(adapter=dspy.JSONAdapter()) 

  • DSPy is integrated with MLFlow (an observability tool)
Reference

Article
https://miptgirl.medium.com/programming-not-prompting-a-hands-on-guide-to-dspy-04ea2d966e6d
https://github.com/miptgirl/miptgirl_medium/blob/main/dspy_example/nps_topic_modelling.ipynb

https://www.dbreunig.com/2025/06/10/let-the-model-write-the-prompt.html

https://thedataquarry.com/blog/learning-dspy-1-the-power-of-good-abstractions/
https://thedataquarry.com/blog/learning-dspy-2-understanding-the-internals/
https://thedataquarry.com/blog/learning-dspy-3-working-with-optimizers/

https://thenewstack.io/goodbye-manual-prompting-hello-programming-with-dspy/

https://dzone.com/static/csrfAttackDetected.html

Paper
https://arxiv.org/abs/2310.03714

Website
https://dspy.ai/
https://dspy.ai/tutorials/games/

Course
https://www.deeplearning.ai/short-courses/dspy-build-optimize-agentic-apps/

Github
https://github.com/stanfordnlp/dspy

Transformers & Large Language Models - 1 of 9


• Background on NLP and tasks

NLP Tasks

1. Classification

- Sentimental analysis :  

* Examples: Amazon reviews, IMDB critiques, Twitter.

* Many to one RNN example. 

Input: sequence of data

Output: scaler. 

- Intent detection

- Language detection

* One to many RNN example. 

Example: Image Captioning and Topic modeling

Input: single or scaler

output: sequence of data

2. "Multi"-Classification

* Synchronous Many to many RNN example. 

Example: Part of speech tagging and Named entity recognition (NER): Dataset = annotated Reuters newspaper (CONLL-2003, CONLL+)

Input: sequence of data

output: sequence of data

- Dependency parsing

- Constituency parsing

3. Generation

* Asynchronous Many to many RNN example. 

Example : Machine translation: Dataset = WMT'14 Translation quality unit is , Question answering,  Summarization, Speech to text

Input: sequence of data

output: sequence of data

Length is not equal. No one to one mapping. 

This RNN example is now done with transformer, LLM. 

- Text generation

History of LLM

1980 RNN

1997 LSTM (Theoretical Foundation) 

2013 Word2Vec

2014 Sequence to Sequence Learning with NN

2015: "Neural machine translation by jointly learning to align and translate" It introduced attention mechanism. Here sequential nature of processing at encoder and decoder. 

2017: Transformer. Parallel processing. "Attention is all you need". Encoder and decoder both have self attention. 

2018: Transfer learning. "Universal language model fine tuning (ULMFit) for text classification" 

- Introduced language modelling

- Now common model for all usecases

- No need of supervised data

- It is about predicating next word. 

Transformer Language Model

1. BERT by Google (encoder only model)

2. GPT by OpenAI (decoder only model). Then GPT2, GPT3 etc. 

2020s LLM

• Tokenization

1. Arbitrary (n/a)

2. Word (multiple tokens with similar meanings need same embedding, so Word variations not handled)  

3. sub-word : focus on common root. Increase sequence length. Tokenization more complex

4. character level: can correct mis-spelled word & CasINg. Sequence length is much longer. No OOV

• Embeddings

Word (Token) Representation by vector

OHE = One Hot Encoding

cosine similarity 

• Word2vec, RNN, LSTM

1. Word2Vec

It is ANN with proxy-task

1. CBOW: Continuous Bag of Words. You predict the target word 

2. Skip-gram : You take the target word and predict words around it

Word order does not matter

Embeddings is not context aware

Dimension size example 768

Special token to indicate "end of sequence" 

2. RNN Recurrent Neural Network

Connection forms a temporal sequence

H = Hidden state = A = Activation Vector = Context Vector. 

RNN is used for all 3 NLP tasks

1. Classification

2. "Multi"-Classification

3. Generation

RNN is keep forgetting the past. This phenomena is called "vanishing gradient"

Word order matters in RNN

3. LSTM = Long short-term memory

1. hidden state

2. cell state

• Attention mechanism

Attention tries to have a direct link between next word that we are predicting and something from the past. 

"self-attention" is main principle of "Attention is all you need" 2017 paper

"self-attention" = Instead of sequential, let direct connection with all part of text at once. 

Concept of Query, Key and Value

We compare Q to K. How they are similar and then take corresponding value

Softmax converts unnormalized network output into probability for different class such that value is [0,1] and sum is 1. 

Formula – Given a query Q, we want to know which key K the query should pay "attention" to with respect to the associated value V. 

attention = softmax ( Q * K ^ T / Sqrt (dimension of K) ) * V

There are three attention layers

1. Attention layer at encoder to compute embeddings from input

2. Decoder-decoder attention OR self-attention layer in decoder, It is is masked, because it only look at those token that are translated. It determines: what other token of output sentence is useful to predict next token. 

3. cross-attention layer : expressed as function of what is seen in input. Last part of encoder. it is fetch to decoder. 

We have direct link to all token. So order words does not matter. (unlike RNN).  So we have Position Encoding: to inform position of word in sequence. 


BOS Token: Beginning of Sequence. 

EOS Token: End of Sequence


• Transformer architecture

Self-attention is achieved by transformer = encoder and decoder

1. Encoder computes meaningful embedding from input text. We have N such encoders. Input layer generates position aware embedding matrix with size d = model size and length = length of input sequence = n

Encoder projects input sequence on 3 spaces Wk, Wq and Wv. so model learns. 

attention = softmax ( Q * K ^ T / Sqrt (dimension of K) ) * V

Projecting on Wq gives a matrix where each row represents a given query Q. So we get matrix Wo that is project back to original dimension of embedding. 

K^T is each column represents key of each token. 

When we multiple K^T and Wq, Each row represents projection of query over each key and then get probability distribution. 

Now multiple with matrix V 

This is self-attention mechanism. means compute representation of each token as function of other tokens. it is done by attention layer. 

Multi-Head Attention (MHA) means this computation is done in different way. So model can learn 

- different representation

- different projections

so all token of input text attend each other. 

It is masked self-attention layer. 

A Multi-Head Attention (MHA) layer performs attention computations across multiple heads, then projects the result in the output space.

Variations of MHA
* Grouped-Query Attention (GQA) and 
* Multi-Query Attention (MQA) 
that reduce computational overhead by sharing keys and values across attention heads.

Head is term given to project matrix that we used to obtain Q, K, V. With more heads, model learns different projection. It is like multiple filters in convolution layer in computer vision. 
h = number of heads
For having h number of heads, the output of attention is h such matrices. Here, because of gradient decent every time we get different result. Each objective function with degree of freedom. We concatenate output of all headers with respect to columns. 

2. FFNN (Feed Forward Neural Network) : so model learn another kind of projection

so we get rich representation of input token

In LLM, hidden layer has higher dimension. So model has enough degree of freedom to learn useful representation. 

3. output is for decoder

It takes Q from output. 

K, V from encoder. 

we have N decoders. 

New Terms

  • Perplexity is an evaluation matrix for machine translation. It quantifies how 'surprised' the model is to see some words together. Lower is better. 
  • OOV = out of vocabulary
  • RNN is keep forgetting the past. This phenomena is called "vanishing gradient"
  • Label Smoothing Purpose

    - prevent overfitting

    - introduce noise

    - let model be little unsure about prediction. 

    It improves accuracy and BLEU score of translation.

  • RLHF : Reinforcement learning from human feedback

References

https://cme295.stanford.edu/

Syllabus : https://cme295.stanford.edu/syllabus/

CheatSheet 

https://cme295.stanford.edu/cheatsheet/ 

https://github.com/afshinea/stanford-cme-295-transformers-large-language-models/tree/main/en 

https://www.youtube.com/watch?v=Ub3GoFaUcds

https://www.youtube.com/watch?v=8fX3rOjTloc

Text Book Super Study Guides

------------------------------------------------------

Sequence to Sequence model has

1. Encoder, Decoder

2. Attention Mechanism

3. Transformer architecture

4. Fine tuning of Transformer Architecture

Usecases

1. Language, sentence has words in sequecne

2. Time series data

3. Biology: Genes, DNA

4. 

------------------------------------------------------

Some more relevant stuff: 

Each layer has 

1. Attention and 

2. Fast Forward


Between two layers we have high dimension 'hidden state vector' in activation space. 


LLM encodes concepts as distributed patterns accross layers = Superposition. 

Antropic has series of papers on superposition and monosemanticity

https://www.youtube.com/watch?v=F2jd5WuT-zg

https://www.neuronpedia.org

https://huggingface.co/collections/dlouapre/sparse-auto-encoders-saes-for-mechanistic-interpretability

https://huggingface.co/spaces/dlouapre/eiffel-tower-llama

------------------------------------------------------------

BAPS IT Convention


On January 18, 2026 BAPS Banglore temple hosted IT convention event from morning to evening. More than 375 participants. 


Here are few take away points 



1.  "Changing Trends of AI in technologies" by Prof. Rahul De

Prof. Rahul De' is founder and CEO of https://www.memoricai.in/ He provided nice academic insight, history of AI, present state and future. AI is about inference and inference is predications, classification and generative output

Human brain has 80 to 86 billion neurons (cells). 

Evolution of AI

1. ANI : Artificial Narrow Intelligence 

2. GNI : General Narrow Intelligence 

3. ANI : Super Narrow Intelligence 

In 1966 a professor Joseph Weizenbaum at MIT developed first chatbot by name ELIZA. It acts like a Rogerian psychotherapist. People like it so much. Later on, we had to convince, that it is not a real person. It is just a computer program that simulates human conversation, through pattern matching and keyword substitution.

Late in 1980 John Searle did "Chinese Room Experiment". Here, a non-Chinese speaker in a room, just manipulate Chinese symbols manually and produce fluent responses without understanding the language.  It proves that just through syntax (rule-following) alone, computer cannot achieve semantics (genuine understanding or consciousness). 

Probably that is why today, GenAI has caveats like hallucinations, jail breaking, bias, privacy violations and unfair responses. In the context of bias, he mentioned about recent movie "Human in the loop"  Available on Netflix. "An indigenous woman works as an AI data-labeler after returning to her village with her children, but soon questions the human bias in machine learning."

He shared some statistics

  • - 2.5 billion prompts are handled by ChatGPT alone in a year. 
  • - 2.4 million models are present at hugging face
  • - 50 billion USD are spent for AI in year 2025. 
  • - 95% firms fail in GenAI adoption. 
  • - We achieved 15% improvement by GenAI

The above numbers raised serious questions that does spending behind GenAI is worth? 

He mentioned few books and categorised all AI adopters in four groups. Boomers (books by Ray Kurzweil), Doomers , Skeptics and Critics "Empire of AI: Dreams and Nightmares in Sam Altman's OpenAI"



2. panel discussion on "Ways to overcome Challenges in IT"


One of the discussion point was about reducing 95% failure rate in 2026

We need: 

  • - Automated task workflow
  • - Cross functional aggregation across various departments
  • - Collaboration between AI and human
  • - Orchestration for AI in day-today work

How can you fail? Your effort (to adapt AI, to retain job etc) can fail. In fact the definition for failure came after industry revolution. During the last century, "Productivity" was more in focus due to industry revolution. 

Other points: 

Former OpenAI co-founder Ilya Sutskever has indicated that simply scaling up to 1 trillion parameters model will not improve AI capability further. 

May be, personalised AI will be next thing

We are humans so 

  • - we are always optimistic, we have hope. 
  • - We have ability to adapt.
  • - We are creator of AI so we are smarter than AI 

We shall remove the fear that we need to learn everything. Yes, we shall learn something new everyday and take its now on NotebookLM. Internet is flooded with many buzzword about AI. We need to separate signal from noise. 

Now learning is not same as degree earning. 

Now, we need to be aware about all domain. That knowledge shall not be gain by asking ChatGPT. 

Few points were discussed about parenting: We shall have 30 min of productive arguments with kids. We will learn AI from the end-user. We shall enable parental control for Internet, OTT. While using AI, be skeptic. You are interacting with product. So product has market, company / organisation behind it, that want to earn profit. AI product is not your friend. More we use LLM, that much brain is unused and lost.  

We need to learn basic values like

  • - Humans are not means
  • - Respect life in people

About firing due to AI. 

  • * If software engineers consider themselves as coder then AI will replace them. They shall consider themselves as problem solver. 
  • * On lighter notes: Pujya Aksharatit Swami mentioned that we SADHUs are easier to get replace. Chat with AI is available round the clock. 
  • * On lighter notes: Pujya Aksharatit Swami mentioned 

येषाम् न विद्या न तपः न दानम् न ज्ञानम् न शीलम् न गुणः न धर्मः ।।

ते मृत्यु-लोके भुवि भार-भूता मनुष्यरूपेण मृगाः चरन्ति ।

It means: Those who possess no knowledge (Vidya), no penance (Tapa), no charity (Dana), no wisdom (Jnana), no good character (Sheela), no virtues (Guna), and no righteousness (Dharma), are a burden to the earth. Although they look like humans, they roam the earth like animals in human form. 

Now this sloka is applicable for knowledge of AI also :-) 

During the panel discussion, the floor is open for everyone to ask question via WhatsApp group, that was flooded with many questions. 

3. Networking

The audience was divided into many groups. The participants had a round of introductions within group. There was engaging quiz, where all groups participated and the group leader responded to questions on behalf of group. 

We had delicious vegetarian, SAATTVIK food. Now some key take away points from post lunch sessions.



4. Work-Life Integration & Wellbeing

“In God we trust. All others must bring data.” - W. Edwards Deming

Here are some shocking data points

  • 83% of Indian IT professional are burnout
  • 73% of European IT professional are burnout
  • 72% are working beyond limit

The data source is ISACA (Information Systems Audit and Control Association)

5 pillars of work life integration

1. Know your why?

What is purpose? It brings impact, mastery and autonomy. 

Our health and family should tune to work. 

2. Design your system

Define boundaries so you can protect capacity

Bring rhythms by frequent breaks. 

3. Recovery and build resilience practices

3.1 Mindful ness and reflection. Moment to moment non-reactive, non-judgmental awareness is mindful ness. 

3.2 physical movement

3.3 social connections

4. Align your environment

5. Seek help early

Something about sleep

Sleep is non-negotiable.

- Sleep hygiene and nutrition  

- Stop all screens 60 to 90 min before going to bed. It is digital sunset. 

- use eye mask while sleeping. 

Next topic was NSANE

Nutrition

Screen Time

Automatic health. He mentioned about book : "Atomic Habit

Notice Signal

Engage with real people

Remember 3 truths

1. Mind = body. Means if body is unhealthy, means mind is unhealthy. Body can be healthy by making healthy mind

2. Work family balance is important

3. It is still not too late. 

Ask your wife, what she expect about husband? 

- Rich husband?

- Healthy husband? 

- Rich but unhealthy husband? 

- burn out person as husband? 

3. Personal/Spiritual Fire Chat with Pujya Santo

Along with other relevant questions and guidance by Pujya SAINTs, again IT layoff was discussed. The apparent reason is AI for layoff. The real reason can be poor performance of employee. Extra hiring happens during COVID phase, so now layoff is inevitable. जातस्य हि ध्रुवो मृत्युः Same way, if you have job, you may get fire. If not today then in future, at age of 62 years. Even after 10 years, the present software application has no values. This is also as per SANKHYA philosophy. It inspires us to make better documentation of product. 



Summary

1. Be happy

2. Worship God. 


બી.એ.પી.એસ. પ્રકાશ એપ


પ્રકાશ વિશેષ :

બેંગલુરુમાં આઇ.ટી. પ્રોફેશનલ્સ માટે યોજાયો વિશેષ સમારોહ


વધુ માહિતી માટે નીચેની લિંક પર ક્લિક કરો.


https://bapsprakash.in/deeplink?data=QsfHnR6PaUp%2FlihVajHZKg%3D%3D%3AWm7u3d6zy%2FylpRGzMz6xFENQBWhBhVioEwCNJCh3MI8%3D


Disclaimer: The author had put best effort to capture all the points, as per his understanding. It may or may not reflect exact intention of the speaker. So any corrections are welcome. This article is not verbatim 

LLMOps


For AI application, we need automation of 

1. Data preparation
2. model tuning
3. Deployment
4. Maintenance and 
5. Monitoring

  • Managing Dependency adds complexity. 

E2E workflow for LLM based application. 

MLOps framework

1. data ingestion

2. data validation

3. data transformation

4. model

5. model analysis

6. serving model

7. logging. 

LLM System Design

boarder design of E2E app including front end, back end, data engineering etc. 

Chain multiple LLMs together

* Grounding : provides additional information/fact with prompt to LLM. 

* Track History. how it works past. 

LLM App

User input->Preprocessing->grounding->prompt goes to LLM model->LLM Response->Grounding->Post processing + Responsible AI->Final output to user.

Model Customization

1. Data Prep

2. Model Tuning

3. Evaluate

It is iterative process

LLMOps Pipeline (Simplified)

1. Data Preparation and versioning (for training data)

2. Supervised tuning (pipeline) 

3. Artifact = config and workflow : are generated. 

- config = config for workflow

E.g. 

Which data set to use

- Workflow = steps 

4. Pipeline execution

5. deploy LLM 

6. Prompting and predictions

7. Responsible AI

Orchestration = 1 + 2 . Orchestration : What is first, then next step and further next step. sequence of step assurance. 

Automation = 4 + 5

Fine Tuning Data Model using Instructions (Hint)

1. rules

2. step by step

3. procedure

4. example

File formats

1. JSONL: JSON Line. Human readable. For small and medium size dataset. 

2. TFRecord 

3. Parquet for large and complex dataset. 

MLOps Workflow for LLM

1. Apache Airflow

2. KubeFlow

DSL = Domain Specific Language

Decorator 

@dls.component

@dls.pipeline

Next compiler will generate YAML file for pipeline

YAML file has

- components

- deploymentSpec

Pipeline can be run on

- K8s

- Vertex AI pipeline execute pipeline in serverless enviornment

PipelineJob takes inputs

1. Template Path: pipline.yaml

2. Display name

3. Parameters

4. Location: Data center

5. pipeline root: temp file location

Open Source Pipeline

https://us-kfp.pkg.dev/ml-pipeline/large-language-model-pipelines/tune-large-model/v2.0.0

Deployment

Batch and REST

1. Batch. E.g. customer review. Not real time. 

2. REST API e.g. chat. More like teal time library. 

* pprint is library to format 

LLM provides output and 'safetyAttributes'

- blocked

* We can find citation also from output of LLM

===========

vertexAI SDK

https://cloud.google.com/vertex-ai

BigQuery 

https://cloud.google.com/bigquery

sklearn

To decide data 80-20% for training and evaluation. 

Building AI/ML apps in Python with BigQuery DataFrames | Google Cloud Blog

===========

AI Bootcamp for students


 8 Day Live Online Workshop

AI Bootcamp for Students

Make Your Child Future-Ready with AI

by Timesof Inida


https://www.notion.com/product Documentation

https://www.todoist.com/ To Do List

https://gamma.app/ For presentation 

https://openai.com/index/sora/ Cinematic Video 

https://www.midjourney.com/home Art Grade Visuals for story telling

https://ideogram.ai/t/explore Typography to image. Communicate in style

https://lovable.dev/ No code web apps

https://n8n.io/ Workflow automation tools

Few more tools

TachyonGPT accelerate the project planning process, potentially saving weeks of effort. This powerful AI assistant allows you to create a complex backlog structure for your project in very little time. Tachyon GPT gives you the power to improve existing work items or generate new work items based on brief titles or descriptions. https://marketplace.visualstudio.com/items?itemName=Neudesic.TachyonGPT

windserf editor and cascade. Agentic code IDE


Reference: https://economictimes.indiatimes.com/masterclass/ai-for-students

https://www.msn.com/en-in/money/news/chatgpt-to-google-gemini-top-5-ai-tools-to-enhance-productivity-mostly-free/ar-AA1GRlt1


Regional LLM, SLM, TinyLM Language Learning


New Language Learning

Want to learn a new language this summer? Explore these expert-led platforms

Mobile app https://youtu.be/jyffkeM9GB0

Regional LLM

Sarvam AI launches Bulbul-v2, its voice model with support for 11 Indian languages

https://asr.iitm.ac.in/models/ 

BharatGen  https://bharatgen.tech/  Bharatgen: First Indigenous Language Ai Model Launched In India News In Hindi - Amar Ujala Hindi News Live - Bharatgen:भारत में लॉन्च हुआ पहला स्वदेशी भाषा Ai मॉडल, 22 भाषाओं में करेगा अनुवाद; दूर होंगी संवाद चुनौतियां and Google to collaborate with IIT Bombay’s BharatGen to build indigenous Indic language model

AIKosha

https://aikosha.indiaai.gov.in/home It looks like huggingface website for India


https://aikosha.indiaai.gov.in/home/toolkit having a list of popular AI related tools.

Kannada Models

nomic-embed-text-v2-moe 
snowflake-arctic-embed2
OpenAIs text-embedding-3-large
Vovage
Cohere
intfloat/multilingual-e5-large-instruct
paraphrase-multilingual
BGE-M3 is based on the XLM-RoBERTa

NVIDIA 
https://build.nvidia.com/openai/whisper-large-v3
https://blogs.nvidia.com/blog/india-ai-mission-infrastructure-models/


BharatGen, 
are building frontier multilingual models using NVIDIA Nemotron and NeMo

Sarvam.ai is open-sourcing its Sarvam-3 series across 22 Indic languages (3B to 100B parameters). 

BharatGen launched a 17B-parameter mixture-of-experts model for public services, agriculture, and security.

CoRover.ai handles 10,000 concurrent users for Indian Railways with 5,000+ daily ticket bookings

Gnani.ai processes 10M+ calls daily across telecom, banking, and hospitality — 15x reduction in inference costs

NPCI is exploring AI for India's UPI payments (something I use every single day at Rameshwaram Cafe and everywhere else in Bengaluru!)

Commotion (backed by Tata Communications) built an AI OS for enterprise workflow automation


TinyML








LLM on mobile



Google for Education

Google introduces Gemini tool for students and educators: How this AI tool will transform classroom teaching