AI & ML interests

None defined yet.

Recent Activity

GoktugD  updated a collection 20 days ago
Werea KVKK PrivacyOps
GoktugD  updated a model 20 days ago
Werea-co/Werea-KVKK-Agent-4B
GoktugD  published a model 20 days ago
Werea-co/Werea-KVKK-Agent-4B
View all activity

GoktugD 
posted an update about 13 hours ago
view post
Post
35
🇹🇷 **You tried to break our Turkish NER model.**

And that's exactly what we wanted.

After sharing Werea-TR-NER, we asked the Hugging Face community to challenge our ~110M parameter model with difficult Turkish sentences.

Some worked.

Some exposed weaknesses.

And that's more valuable than pretending a benchmark score tells the whole story.

Our published WikiANN-tr result:

**91.7% Entity F1**

👤 PERSON → 94.2%
📍 LOCATION → 91.4%
🏢 ORGANIZATION → 89.2%

But now we want to go further.

🔥 **ROUND 2**

Send me a Turkish sentence designed specifically to break the model.

Ambiguous names.
Companies that sound like people.
Locations hidden inside organization names.
Turkish suffixes.
Slang.
Anything nasty.

**Try to make it fail.**

I'll collect the hardest examples and use them to build a public adversarial Turkish NER evaluation set.

🤗 Model:
Werea-co/Werea-TR-NER

🇹🇷 Werea:
Werea-co


Follow me if you want to see whether the community can break it — and what we build from the failures.

#TurkishNLP #HuggingFace #NER #OpenSourceAI
GoktugD 
posted an update 3 days ago
view post
Post
2054
🇹🇷 **Can a 110M model understand Turkish names, places and organizations this well?**

We tested Werea-TR-NER on the human-labeled WikiANN Turkish test set:

**91.7% Entity F1**

👤 Person → **94.2%**
📍 Location → **91.4%**
🏢 Organization → **89.2%**

Only ~110M parameters.

Try it with a difficult Turkish sentence 👇

Ahmet Yılmaz İstanbul'da Werea şirketinde çalışıyor.

→ Ahmet Yılmaz — PERSON
→ İstanbul — LOCATION
→ Werea — ORGANIZATION

But easy examples are boring.

**Give me the hardest Turkish sentence you can think of.**

I'll run the most interesting ones through the model and share the failures too.

🤗 Model:
Werea-co/Werea-TR-NER

🇹🇷 Werea:
Werea-co


**Follow Werea if you're interested in open Turkish AI — we're publishing the models, benchmarks and failures openly.**

#TurkishNLP #HuggingFace #NER #OpenSourceAI
GoktugD 
posted an update 7 days ago
view post
Post
2432
🇹🇷 One of our small Turkish models quietly reached **500+ monthly downloads** on Hugging Face.

**Werea-TR-TextRestore — only 300M parameters.**

Its job is simple:

istanbulda hava cok guzel
İstanbul'da hava çok güzel.

A lightweight model for restoring Turkish text:
• diacritics
• punctuation
• casing
• corrupted text

**96.5% word accuracy** on real Turkish news sentences.

And it runs without sending your text to a cloud API.

🤗 Try the model:
Werea-co/Werea-TR-TextRestore

🇹🇷 Built in Türkiye. Open source.

If you're working on Turkish NLP, I'd love to hear what we should build next.

#TurkishNLP #HuggingFace #OpenSourceAI #NLP
  • 2 replies
·
GoktugD 
posted an update 15 days ago
view post
Post
1916
🇹🇷 We started with one question:

**How much of the Turkish AI stack can we build openly?**

Today, Werea has grown to **19 open models on Hugging Face.**

Not just LLMs.

📄 Document AI — Werea-DocOCR-1B
🛡️ Cybersecurity — Werea-NanoSOC-8B
🔐 Privacy / KVKK — Werea-KVKK-Agent-4B
🔍 Retrieval — DUSUNEN-Rota-270M
🎙️ Speech — Werea-TSS
🧠 Turkish NLP — NER, NLI, QA, Intent, Sentiment, Topic & more
👁️ Computer Vision — Gemstone

And we want the results to be measurable.

Some of our published benchmarks:

🏷️ NER → **91.7% F1** — WikiANN-tr
🗂️ Topic → **92.8% accuracy** — TTC4900
🎯 Intent → **88.2% accuracy** — MASSIVE-tr
🧩 NLI → **74.5% accuracy** — XNLI-tr
❓ QA → **72.7% F1** — TQuAD2
📄 DocOCR → **0.15% CER** on our held-out Turkish enterprise document test

Our goal isn't to upload as many models as possible.

Our goal is to build an **open Turkish AI ecosystem**:

Models.
Datasets.
Benchmarks.
Demos.
Real applications.

Built from Türkiye. 🇹🇷
Open to everyone.

🤗 Explore Werea:
Werea-co


📄 Document AI:
Werea-co/Werea-DocOCR-1B

🛡️ NanoSOC:
Werea-co/Werea-NanoSOC-8B

🔐 KVKK Agent:
Werea-co/Werea-KVKK-Agent-4B

🔍 Rota:
Werea-co/DUSUNEN-Rota-270M-v3

If you're building Turkish AI, follow Werea — there's much more coming.

GoktugD 
posted an update 19 days ago
view post
Post
2372
🇹🇷 We trained a 1B OCR model specifically for Turkish enterprise documents.

**Werea-DocOCR-1B v2**

The result surprised us:

LightOnOCR-2 base → **64.2% CER**
Werea-DocOCR v1 → **~8.1% CER**
Werea-DocOCR v2 → **0.15% CER** 🚀

Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.

📄 12 Turkish enterprise document types
🧪 12,960 synthetic training pages
📱 Digital + scanned + phone photos
📊 Tables → structured Markdown
⚙️ Full-parameter fine-tuning
🖥️ Trained on a single RTX 3090

It handles:

• e-Invoices
• rental contracts
• bank receipts
• payroll documents
• insurance policies
• vehicle documents
• official correspondence
• trade registry documents
• SGK-style tables
• and more.

**Model 🤗**
Werea-co/Werea-DocOCR-1B

**Dataset 📚**
Werea-co/werea-tr-doc-ocr-enterprise-v2

**Werea 🇹🇷**
Werea-co


We're building open AI models from Türkiye.

This is just the beginning.

#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision
girenit 
updated a Space 21 days ago