alifaraz commited on
Commit
e21a8b2
·
verified ·
1 Parent(s): 866ef7d

Added usage instructions

Browse files
Files changed (1) hide show
  1. README.md +95 -3
README.md CHANGED
@@ -50,6 +50,98 @@ metrics:
50
 
51
  ---
52
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
53
  ## Evaluation Results
54
 
55
  ### **Indic OCR Performance**
@@ -117,9 +209,9 @@ Note that the above latency values are calculated assuming the size of the input
117
 
118
  ### Observations
119
 
120
- - 🕒 **TTFT (Time-to-First-Token):** ~125 ms
121
- - ⚙️ **Inter-token latency:** ~4 ms per token
122
- - 🧮 **Language impact:** Latency varies with tokenization efficiency.
123
  - **English** and **Hindi** → Lower latency due to **compact token-to-word ratios**.
124
  - **Telugu** and **Malayalam** → Higher latency due to **fragmented tokenization** (larger number of tokens per word).
125
 
 
50
 
51
  ---
52
 
53
+ ## Usage
54
+ ### Using transformers
55
+ ```python
56
+ from PIL import Image
57
+ from transformers import AutoTokenizer, AutoProcessor, AutoModelForImageTextToText
58
+
59
+ model_path = "krutrim-ai-labs/Chitrapathak-2"
60
+
61
+ model = AutoModelForImageTextToText.from_pretrained(
62
+ model_path,
63
+ torch_dtype="auto",
64
+ device_map="auto",
65
+ attn_implementation="flash_attention_2"
66
+ )
67
+ model.eval()
68
+
69
+ tokenizer = AutoTokenizer.from_pretrained(model_path)
70
+ processor = AutoProcessor.from_pretrained(model_path)
71
+
72
+
73
+ def perform_ocr(image_path, model, processor, max_new_tokens=4096):
74
+ image = Image.open(image_path)
75
+ messages = [
76
+ {"role": "system", "content": "You are a helpful assistant."},
77
+ {"role": "user", "content": [
78
+ {"type": "image", "image": f"file://{image_path}"},
79
+ {"type": "text", "text": "Perform OCR on this image and transcribe all visible text exactly as it appears."},
80
+ ]},
81
+ ]
82
+ text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
83
+ inputs = processor(text=[text], images=[image], padding=True, return_tensors="pt")
84
+ inputs = inputs.to(model.device)
85
+
86
+ output_ids = model.generate(**inputs, max_new_tokens=max_new_tokens, do_sample=False)
87
+ generated_ids = [output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids)]
88
+
89
+ output_text = processor.batch_decode(generated_ids, skip_special_tokens=True, clean_up_tokenization_spaces=True)
90
+ return output_text[0]
91
+
92
+ image_path = "/path/to/your/document.jpg"
93
+ result = perform_ocr(image_path, model, processor, max_new_tokens=15000)
94
+ print(result)
95
+ ```
96
+
97
+ ### Using vLLM
98
+ 1. Start the vLLM server.
99
+ ```bash
100
+ vllm serve krutrim-ai-labs/Chitrapathak-2
101
+ ```
102
+ 2. Predict with the model
103
+ ```python
104
+ from openai import OpenAI
105
+ import base64
106
+
107
+ client = OpenAI(api_key="123", base_url="http://localhost:8000/v1")
108
+
109
+ model = "krutrim-ai-labs/Chitrapathak-2"
110
+
111
+ def encode_image(image_path):
112
+ with open(image_path, "rb") as image_file:
113
+ return base64.b64encode(image_file.read()).decode("utf-8")
114
+
115
+ def perform_ocr(img_base64):
116
+ response = client.chat.completions.create(
117
+ model=model,
118
+ messages=[
119
+ {
120
+ "role": "user",
121
+ "content": [
122
+ {
123
+ "type": "image_url",
124
+ "image_url": {"url": f"data:image/png;base64,{img_base64}"},
125
+ },
126
+ {
127
+ "type": "text",
128
+ "text": "Perform OCR on this image and transcribe all visible text exactly as it appears.",
129
+ },
130
+ ],
131
+ }
132
+ ],
133
+ temperature=0.0,
134
+ max_tokens=15000
135
+ )
136
+ return response.choices[0].message.content
137
+
138
+ test_img_path = "/path/to/your/document.jpg"
139
+ img_base64 = encode_image(test_img_path)
140
+ print(perform_ocr(img_base64))
141
+ ```
142
+
143
+ ---
144
+
145
  ## Evaluation Results
146
 
147
  ### **Indic OCR Performance**
 
209
 
210
  ### Observations
211
 
212
+ - **TTFT (Time-to-First-Token):** ~125 ms
213
+ - **Inter-token latency:** ~4 ms per token
214
+ - **Language impact:** Latency varies with tokenization efficiency.
215
  - **English** and **Hindi** → Lower latency due to **compact token-to-word ratios**.
216
  - **Telugu** and **Malayalam** → Higher latency due to **fragmented tokenization** (larger number of tokens per word).
217