What GenAI? Part 1


What GenAI?

GenAI is about generating text, images, or other media, in response to prompts. GenAI is using generative models. GenAI is based on Transformer based deep learning model

Modality

- Unimodal (only 1 input)

- Multimodal. E.g. GPT-4 accepts text and image. Wu Dao

================================================

What Transformer?

* A transformer is a deep learning architecture, 

* It is designed to understand the context and semantics of language.

* It takes in a sequence of tokens (words, or parts of words) and outputting a corresponding sequence. 

* It pays attention to each input token and the relationships between them, using a mechanism known as 

1. self-attention or 

2. scaled dot-product attention. 

* This enables transformer to understand complex linguistic constructs and generate coherent and contextually accurate responses.

* It relies on the parallel multi-head attention mechanism.

* requiring less training time than previous recurrent neural architectures, (e.g. long short-term memory (LSTM) )

Implementation

- TensorFlow

- PyTorch

- JAX Deep Learning

- Transformer library by Hugging Face

Architecture

1. Tokenizer

2. Embedding layer. Token to vector

3. Transformer Layers : alternate attention and feedforward. 

4. Optional un-embedding layer

* It uses activation function ReLU, SwiGLU

================================================

What ChatGPT?


It is LLM chatbots. It enables users to refine and steer a conversation towards a desired length, format, style, level of detail, and language. 

ChatGPT is based on GPT 3.5 and ChatGPT Plus is based on GTP 4

Other examples: Bing Chat (based on GTP 4), Bard, LLaMA, Ernie Bot

Usage: 
  • write and debug computer programs
  • teleplays
  • fairy tales
  • student essays
  • answer test questions 
  • generate business ideas
  • write poetry 
  • song lyrics
  • translate and summarize text
  • emulate a Linux system
  • simulate entire chat rooms
  • play games like tic-tac-toe
  • or simulate an ATM.
================================================
What GPT?

generative pre-trained transformers
Based on uni-directional ("autoregressive") transformers
It is a GenAI model. It combines two forms of training

1. Pre-Training: General purpose, using vast quantities of data
2. Fine-Tuning: Supervised ML tasks on small specific data. 

================================================
What OpenAI?

an organization behind GPT and other products
================================================
What LLM?

Language Model (LM) It is ML approach to model, probability distribution over a sequence of words. It predicts probability for next word in a sequence. 

LLM is large ANN with billions of parameters. It is trained on large quantities of data using self-supervised / semi-supervised approaches. LLM is GenAI for language / text. 

Examples: GPT-2, GPT-3, GPT-4, GPT-J, Claude, BERT, XLNet, RoBERTa, BLOOM (BigScience Large Open-science Open-access Multilingual Language Model), LaMDA, LLaMA, Stable Diffusion, PaLM, FLAT T-5, Llama, gpt4all, Llama 2, Code Llama, Mistral

Smaller models
- LLaMA-7B (Raspberry Pi 4)
- one version of Stable Diffusion on iPhone 11
- Llama-2

Deploy

OctoML allows to host model on server and edge devices (even on browser)
================================================
What GenAI Stack?

1. Data Extraction and loading (airbyte and llamahub) 
2. Embeddings (Word2Vec, GloVe, and FastText) 
3. Vector DB
4. Prompt Engine
5. Retrieval
6. Memory
7. Model

================================================

What DGM Deep Generative Model ?

A generative model is a statistical model of the joint probability distribution P(X,Y) on given observable variable X and target variable Y.

Simple example

Suppose the input data is , the set of labels for  is , and there are the following 4 data points: 

For the above data, estimating the joint probability distribution  from the empirical measure will be the following:



Types

1. variational autoencoders (VAEs)
2. generative adversarial networks (GANs)
3. auto-regressive (uni-directional) models

Examples
For text
1. GPT2
2. GPT3
3. Bidirectional Encoder Representations from Transformers (BERT)
For image
1. BigGAN
2. VQ-VAE
        3. DALL-E
For Music
1. Jukebox
        2. MuseNet
        3. MusicLM
        4. MusicGen
For Text to Video
        1. RunwayML
        2. Make-A-Video by Meta Platforms
For Programming
        1. GitHub Copilot
For text to image
        1. Midjourney
    

This architecture has also led to the development of pre-trained systems, such as generative pre-trained transformers (GPTs) and BERT[12] (Bidirectional Encoder Representations from Transformers).

================================================

What Vector DB ?

Word embedding : encode each word from training set as vector. It is a representation of a word. It encodes the meaning of the word in such a way that words that are closer in the vector space are expected to be similar in meaning. It is useful for syntactic parsing and sentiment analysis.

The OpenAI word embedding model lets you take any string of text (up to a ~8,000 word length limit) and turn that into a vector with 1536 dimension. So word has 1,536 floating point numbers as attributes. These floating point numbers are derived from a sophisticated language model. They take a vast amount of knowledge of human language and flatten that down to a list of floating point numbers. 4 bytes per floating point number that’s 4*1,536 = 6,144 bytes per word embedding—6KiB.

the whole vocabulary is vector DB. It is useful for sequence prediction.

Example: Pinecone, Weavite, chromadb, drant, activeloop, pgvector, zilliz, redis, momento, Neo4j, Casandara (CaasIO library)

================================================

What NLP Tasks?
  • machine translation
  • document summarization
  • document generation
  • named entity recognition (NER)
  • biological sequence analysis
  • writing computer code based on requirements expressed in natural language.
  • video understanding.
  • syntactic parsing
  • sentiment analysis
================================================

What Prompt Engineering?

process of structuring text that can be interpreted and understood by a generative AI model. It is enabled by in-context learning, defined as a model's ability to temporarily learn from prompts.

Example: PromptLayer, Aim, scale, Humanloop, HoneyHive

LangChain and LlamaIndex are useful. 

================================================

What RAG?

retrieval augmented generation

It can finetune the models 
We can feed an initial prompt with additional data from live database. 
It enables to personalize or finetune an answer on the fly.

Retrieval Augmented Generation (RAG) is a method for improving the performance of large language models (LLMs) by providing them with access to external knowledge sources. This is done by first retrieving a set of relevant documents from the knowledge source, and then using those documents to generate a response.


Two sides of RAG

1. Semantic Search
2. Cypher Generation 

Examples

1. ChatGPT plugin
2. Google Search
3. Vector DB
4. Knowledge Graph (graph DB Neo4J)
5. LlmaIndex (GPT Index) is also an library of LangChain
https://gpt-index.readthedocs.io/en/latest/index.html

It has
5.A. Data connectors API, PDF, SQL etc.
5.B. Data Index : Structure data as intermediate representation
5.C. Engine : Natural language access to data e.g. Chat Engine , Query Engine
5.D. Data Agents : LLM powered knowledge worker
5.E. Application Integrations : tie LlamaIndex back into the rest of your ecosystem. This could be LangChain, Flask, Docker, ChatGPT, etc.

RAG and Prompt Engineering are two of the techniques to eliminate Hallucination 

================================================

What Hallucination ?

GenAI generates output that looks very authenticate but actually it is false, untrue, incorrect. Like Fake Video. However Fake Video is not result of hallucination. 

Why Hallucination?

1. The input parameter temperature is high. So chances of output will go off track is higher. 

2. Missing Information. E.g. all LLM models have some cut off date. There is no information about event after that cutoff date. 

3. Bias training data and complex models. 

How to avoid hallucination ?

1. Prompt Engineering
2. In-context learning
3. Fine tuning of actual model
4. Grounding using RAG. RAG / grounding is available since May 2020

Food Items and other useful stuff


अन्नम Rice 

लवण salt 

क्वथितं Sambhar 
व्यंजनं Veg. gravy 
कोषंभरी Veg.  Salad 
वेसवार: Pody  powder 
दधि curd 
तक्रम् buttermilk 
अवलेह: pickle 
पर्पट: Papad 
पायसम: 
कुण्डलिका Jalebi 
​रोटिका 
पुरीका 
====================
प्रातराश: Breakfast 

इडली 
उपसेचन Chatani 
दोसा 
वटक: Vada 
संधित खाद्यं Sandwitch 
सुपिष्टकम् Bread 


====================
क्रिणाति क्रेष्यति 
====================
शिवस्य पुत्र: गणेश: मार्गे स्युतात् फलं हस्तेन मित्राय ददाति 
====================
मुनिवर्य:
स्वामीवर्य: 
आंजनेय:
महोदय:
====================
डिम्ब: कीलानाम् अभ्यवहरति = शिशु: दुग्धं  पिबति 
====================
क्रिड़ालु:
गायक:
भाषणकार:
चालक:
लेखक:
वैद्य:
नर्तकी 
अध्यापिका 
====================
यद्यपि ---- तथापि 

निर्धन: -- दानं 
बुद्धिमान -- अनुर्तीण  
भोजन -- बुभुक्षा 
अधिकं पठति -- स्मरणे न तिष्ठति 
कष्टं -- योगासनम्

Everything about PKI


entity = computer | user

Every entity has identity

Authentication is verification of claim by entity

======================================

Hash function: same input => same output. If input is different by even a bit, output is completely different. They are one way. MAC (Message Authentication Code) is hash function. MAC needs common key. 

Signature is similar but uses key pair.  If only one entity knows the private key you get a property called non-repudiation: the private key holder can't deny (repudiate) the fact that they signed some data.

sign with private key. verify with public key. 

encrypt with public key. decrypt with private key

public key cryptography = asymmetric cryptography, as above

=====================================

subscriber or end entity is subject of certificate

CA is issuer of certificate

CA has root certificate | intermediate certificate

end entity has leaf certificate

relying party trust CA and verify certificate.

relying parties are pre-configured with a list of trusted root certificates (or trust anchors) in a trust store. Root certificates in trust stores are self-signed . Trust store by 4 major organization

1. Apple's root certificate

2. Microsoft's root certificate programm

3. Mozila's root certificate programm

2. Google's root certificate programm

Cloudflare's cfssl project maintains a github repository that includes the trusted certificates from various trust stores. https://github.com/cloudflare/cfssl_trust

=======================================

Certificate : issues says entity (subject) has public key

X.509, ASN.1, OIDs, DER, PEM, PKCS

X.509 builds on ASN.1. ASN.1 : DER (Binary) and BER

PEM: base64 encoded DER payload sandwiched between a header and a footer

Possible extension for PEM: .der, .pem, .crt, .cer

======================================

PKCS (Public Key Cryptography Standards) published by RSA labs

  • * PKCS#7 is rebranded as Cryptographic Message Syntax (CMS) is by IETF. It is used by Java.

Possible extension for PKCS#7: .p7b, .p7c

  • * PKCS#12 = certificate chain + private key

It is used by microsoft products

Possible extension for PKCS#12: .pfx, .p12

So possible formats for PKCS#7 and PKCS#12 are : raw der, pem, ber

raw der is most widely used.

  • * PKCS#8 is for private key and its metadata

Possible extension for private key: .prv, .key, .pem

Possible extension for public key: .pub, .pem

They may includes headers: Proc-Type, DEK-Info

  • * PKCS#10 is CSR

====================================

  • SSH: certificate less PKI .It binds names to public key in files
  • PGP: uses certificate, but not CA. It uses web-of-trust model
  • Web PKI (Internet PKI or PKIX) works with browser.  no control over important details like certificate lifetime, revocation mechanisms, renewal processes, key types, and algorithms

=====================================

Bundle of certificate: root intermediate leaf - forms certificate chain. More often, certificate chains are encoded as a simple sequence of line-separated PEM objects. Some stuff expects the certs to be ordered from leaf to root, other stuff expects root to leaf, and some stuff doesn't care. The relying party verifies the leaf and intermediate certificates in a process called certificate path validation.

=====================================

PKIX originally specified to use FQDN in DN common name. The modern best practices is to leverage SANs. There are four sorts of SANs in common use, all of which bind names that are broadly used and understood: domain names (DNS), email addresses, IP addresses, and URIs.

=====================================

  • To configure a PKI relying party you tell it which root certificates to use
  • To configure a PKI subscriber you tell it which certificate and private key to use (or tell it how to generate its own key pair and exchange a CSR for a certificate itself)


Helm


The Helm Parent folder has 

1. following folders

* templates

* charts

2. following files

* Chart.yaml

* values.yaml

Templates folders has all K8s object. 

Here is list of built-in objects https://helm.sh/docs/chart_template_guide/builtin_objects/

The important ones are: 

1. Release

This object describes the release itself. It has several objects inside of it:

  • Release.Name: The release name
  • Release.Namespace: The namespace to be released into (if the manifest doesn’t override)
  • Release.IsUpgrade: This is set to true if the current operation is an upgrade or rollback.
  • Release.IsInstall: This is set to true if the current operation is an install.
  • Release.Revision: The revision number for this release. On install, this is 1, and it is incremented with each upgrade and rollback.
  • Release.Service: The service that is rendering the present template. On Helm, this is always Helm.

2. Chart

The contents of the Chart.yaml file. The fields are as per https://helm.sh/docs/topics/charts/#the-chartyaml-file

apiVersion: The chart API version (required)
name: The name of the chart (required)
version: A SemVer 2 version (required)
kubeVersion: A SemVer range of compatible Kubernetes versions (optional)
description: A single-sentence description of this project (optional)
type: The type of the chart (optional)
keywords:
  - A list of keywords about this project (optional)
home: The URL of this projects home page (optional)
sources:
  - A list of URLs to source code for this project (optional)
dependencies: # A list of the chart requirements (optional)
  - name: The name of the chart (nginx)
    version: The version of the chart ("1.2.3")
    repository: (optional) The repository URL ("https://example.com/charts") or alias ("@repo-name")
    condition: (optional) A yaml path that resolves to a boolean, used for enabling/disabling charts (e.g. subchart1.enabled )
    tags: # (optional)
      - Tags can be used to group charts for enabling/disabling together
    import-values: # (optional)
      - ImportValues holds the mapping of source values to parent key to be imported. Each item can be a string or pair of child/parent sublist items.
    alias: (optional) Alias to be used for the chart. Useful when you have to add the same chart multiple times
maintainers: # (optional)
  - name: The maintainers name (required for each maintainer)
    email: The maintainers email (optional for each maintainer)
    url: A URL for the maintainer (optional for each maintainer)
icon: A URL to an SVG or PNG image to be used as an icon (optional).
appVersion: The version of the app that this contains (optional). Needn't be SemVer. Quotes recommended.
deprecated: Whether this chart is deprecated (optional, boolean)
annotations:
  example: A list of annotations keyed by name (optional).

3. Values

Values passed into the template from the values.yaml file and from user-supplied files.

Now commands

1. helm lint .

2. helm template .

3. helm install --dry-run my-release

4. helm install my-release

5. helm list

We can pass external values.yaml using "--values /path/values.yaml"

6. helm upgrade my-release

7. helm rollback my-release

8. helm rollback <release-name> <revision-number>

9. helm uninstall my-release

10. helm package my-release

to put it on GitHub, S3 etc. 

Debugging of Helm

1. helm lint

2. helm get values: This command will output the release values installed to the cluster.

3. helm install --dry-run

4. helm get manifest: This command will output the manifests that are running in the cluster.

5. helm diff: It will output the differences between the two revisions.

Reference: https://devopscube.com/create-helm-chart/

Data on Kubernetes


storage system attributes: 

- availability [ # of replica , primary and secondary DB ]

the ability to access the data during failure conditions. The failures may be due to failures in the storage media, transport, controller or any other component in the system.

RTO Recovery Time Objective, MTTF, MTTR

- consistency

a. eventually consistent

b. strongly consistent

RPO Recovery Point Objective (time)

- scalability [ sharding = divide larger part into smaller part ]

a. number of client

b. throughput

c. capacity

d. increase number of component to support all above. 

- durability 

a. [ # of replica ]

b. endurance characteristic of storage media: SSD, spinning disk, tape

c. ability to detect corruption of data and recover corrupted data (bit-rot)

- performance 

a. latency

b. operation per second

c. throughput for read and throughput for write. 

For cloud native storage system : three more

- observability

- elasticity [ on demand scale up/down ] 

- data locality [ pod affinity ]

Storage stacks / layers

1. Data Access Interface [ block device, file system, App API: (Object store, k-v store and DB), PIP

2. Storage topology [ centralized, distributed, sharded, and hyper-converged ]

3. data protection layer, which adds redundancy [ RAID, Erasure coding, and Replicas ]

4. additional data services [ replication, snapshots (PiT Point in Time), clones, incremental snapshots for efficient backups ]

5. host, OS, physical non-volatile storage

Common Patterns and Features

1. Operator

2. CSI

2.1 CSI building blocks

2.1.1 identity gRPC service : info and capabilities of plugin

2.1.2 controller gRPC service: 

2.1.2.1. create and delete volume, 

2.1.2.2. create and delete snapshot, 

2.1.2.3. attach and detach volume, and 

2.1.2.4. expand volume

2.1.3. node gRPC service

2.1.3.1. mount and unmount volume, and 

2.1.3.1. expand volume.

2.2 CSI features

2.2.1 CSI topology and CSI capacity tracking provides input to K8s scheduler. 

2.2.2 raw block mode ( instead of file system)

2.2.3 snapshot and group snapshot for backup / recovery. 

3. K8s workload API

Volume Claim Template

4. Topology Aware Scheduling

5. Pod Disturption Budget

6. Resource Management

6.1 pod's [ guaranteed ] QoS, 

6.2 pod's [ higher than normal ] priority

6.3 VPA is better than HPA for statefulset. 

7. Separation of CP (using operator) and DP (E-W traffic)

8. Default secure

8.1 no port accessible outside

8.2 k8s secret. 

Day 2 operations

Upgrade

- CRD version

Backup/Restore

- App level

- volume level with hook so during backup, app does not use volume 

Increase / Decrease Storage capacity

- stateful set can expand storage volume

- HPA

Data Migration

- init containers

- job

Reference: