Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

SeaWolf-AI 
posted an update 1 day ago
view post
Post
2791
AX-Ray: Safety Diagnostics for AI/AX Models

AI models can no longer be evaluated only by capability scores. As models move into public services, enterprise workflows, scientific research, and administrative decision support, we need a second layer of evaluation: whether the model behaves safely, structurally, and consistently under real deployment conditions.

VIDRAFT AX-Ray is a public AI/AX safety diagnostic initiative powered by FINAL-Bench Diagnostics. AX-Ray evaluates models across a structured guideline framework, including model-level safety, AX deployment readiness, and agent/service operation risks. The public diagnostic catalog contains 117 diagnostic items, mapped to legal, regulatory, ethical, and religious-law governance contexts so that safety review can be discussed in a form closer to real institutional responsibility.

A central finding of AX-Ray is causal leakage: a structural defect where information that should not influence an earlier reasoning state appears to affect model behavior. AX-Ray presents a public case of diagnosing, reproducing, and demonstrating causal leakage in two general-purpose public models. This matters because such defects are not exposed by ordinary benchmark scores. A model can appear capable while still carrying hidden safety or integrity risks.

Explore the live leaderboard, diagnostic reports, and public dataset here:

- AX-Ray Space: FINAL-Bench/AX-RAY
- AX-Ray Dataset: FINAL-Bench/AX-RAY
- Technical Article: https://huggingface.co/blog/FINAL-Bench/ax-ray

AX-Ray is intended as a practical guideline for moving AI evaluation beyond “how smart is the model?” toward “can this model be trusted, governed, and deployed safely?”
  • 2 replies
·
SeaWolf-AI 
posted an update about 16 hours ago
view post
Post
918
🧬 Your AI can design a malaria drug candidate. Can it tell you whether it's any good?

Open Discovery Challenge #1 — Malaria is live. Design a molecule with any model — OpenAI, Claude, Gemini, Qwen, KIMI, DeepSeek, open weights, or by hand — submit it as SMILES, and it's scored in minutes on whole-cell activity, target binding, selectivity over the human enzyme, ADMET, novelty and synthesisability.

You can check the scoring instead of trusting it. Approved drugs sit on the same leaderboard as the entries: DSM265, a clinical-stage antimalarial, scores 50.9. Teriflunomide — approved, but it hits the human enzyme — scores 2.8. Caffeine scores 1.8. If the clinical candidate lands on top and coffee lands at the bottom, the scorer discriminates.

We caught 14 defects before opening — conventional toxicity cutoffs rejected all three approved antimalarials and coffee. All written up, along with the rule we now hold everything to: a gate that rejects an approved drug is a broken gate.

Your molecule stays yours. No patent interest, nothing into our pipeline. You choose whether it's published — and publishing can cost you patentability, so we say so.

USD 1,000 to the top entry when Season #1 closes 30 September 2026 — not payment for your tokens, but a way of saying the work had worth.

Malaria killed ~597,000 people in 2023, three quarters of them children under five. Not for want of chemistry — for want of a market.

No chemistry needed: the guide ships five prompts you can paste straight into your model, and the full rubric is published.

📖 https://huggingface.co/blog/FINAL-Bench/open-discovery-challenge
🚀 FINAL-Bench/open-discovery-challenge

Computational assessments of candidates — not measurements, not claims of efficacy.
Banaxi-Tech 
posted an update 1 day ago
view post
Post
1399
We're excited to release BananaMind 2 Pro, our final version of the Pro model.
Trained on 100B tokens it performs extremely good for its token and size class.
The training took 22 days on one RTX 5070 Ti.
Check it out at
BananaMind/BananaMind-2-Pro
We did not release a Chat version yet because it regressed. Release Later.
Follow us to know when BananaMind 2 Ultra releases and support us at
BananaMind

@Banaxi-Tech
@vovaRL
@DedeProGames
  • 2 replies
·
onekq 
posted an update 2 days ago
view post
Post
2685
This is an easy-to-remember pattern.

GLM and Kimi are in Beijing, DeepSeek and Qwen are in Hangzhou.

OpenAI and Anthropic are in SF, xAI and Meta are in the peninsula.

In both China and the Bay Area, token price of one area is ~1/3 of the other area.
  • 1 reply
·
JonathanColetti 
posted an update about 23 hours ago
view post
Post
1679
I just uncensored qwen 3.8 27B. You can check out the demo here: JonathanColetti/Qwen3.8-27B-Uncensored-Demo or the full repo: JonathanColetti/Qwen3.8-27B-Uncensored-GGUF

What I did was used an H100 NVL and use https://github.com/p-e-w/heretic to uncensor it. I measured the KL divergence and the refusal rate (Refusal rate is the count of refusals over 100 hold out prompts from mlabonne/harmful_behaviors). Check out the full repo and give it a like if you think its cool

Thanks!
YatharthS 
posted an update about 20 hours ago
view post
Post
1065
Open sourcing Coala Embeddings! It can compress 20 million documents which can take 80gb ram down to just 1gb RAM.

Why is this important? Embeddings are used everywhere(RAG, semantic search, etc.) by many.

Coala Embeddings massively cheapen cost to host them while being simply and easy to use.

GitHub: https://github.com/ysharma3501/CoalaEmbeddings
Notebook Demo: https://colab.research.google.com/drive/1TXip-vZg72e_2AB185MINufVVav8Oo9T?usp=sharing
HannesVonEssen 
posted an update 1 day ago
view post
Post
642
👀 Qwen3.8-27B is identical to Qwen3.6-27B!

Interestingly, the just released 3.8 version has exactly the same architecture - meaning all the capability gains come from training improvements!

See the diff (0 changes) here!

https://hfviewer.com/compare/qwen3.6-27b-vs-qwen3.8-27b
  • 1 reply
·
sequelbox 
posted an update 2 days ago
view post
Post
2899
NEW RELEASES for the new Muse Glimmer 30B!

- Esper 4, our flagship agentic coder: specialist in coding, architecture, DevOps, and MLOps!
- Tachibana-Agent, trained only on code for dedicated, predictable deployment!

GET OUR NEW MODELS:
ValiantLabs/Muse-Glimmer-30B-Esper4
sequelbox/Muse-Glimmer-30B-Tachibana-Agent

Get the datasets for your own training:
sequelbox/Titanium4-DeepSeek-V4-Pro
sequelbox/Mitakihara2-DeepSeek-V4-Pro
sequelbox/Tachibana4-DeepSeek-V4-Pro

We'll be expanding Esper 4 to more models and releasing new models as funding allows - donate for more, faster, better models and datasets: sequelbox/SupportOpenSource

Also, starting work on the SV4 datasets :)

More to come soon!

love,
allegra
  • 1 reply
·
ezgikorkmaz 
posted an update 1 day ago
Javedalam 
posted an update 1 day ago
view post
Post
643
A 3B Vision Model Running at 13 Tokens/Second — on a Phone

AI is moving from the cloud into your pocket.

I just tested Liquid AI's newly released LFM2.5-VL-3B locally on a Samsung Galaxy S26. The Q8_0 model ran through llama.cpp with Vulkan acceleration on the Qualcomm Adreno 840 GPU. I uploaded a real webpage screenshot from another device and asked the model to describe it.

It worked—and generated at 13.11 tokens per second.

That's significant for a 3B-class vision-language model running entirely on a phone. In my testing, LFM2.5-VL-3B is the most capable local vision model I've run on the S26 so far.

Liquid AI is building specifically for this world. Its focus is edge AI: capable models designed to run locally on phones, laptops, vehicles and other resource-constrained hardware instead of depending entirely on cloud inference. Its LFM family targets low memory usage, fast inference and deployment across CPUs, GPUs and NPUs.

LFM2.5-VL-3B brings vision into that strategy. It is an open-weight multimodal model capable of image understanding, OCR, document extraction, visual question answering and spatial reasoning. Liquid provides the weights publicly, including an official GGUF release for llama.cpp.

This is where open-weight AI gets interesting. A phone can now hold the model, process the image and generate useful multimodal responses locally at interactive speed.

No cloud GPU. No remote inference API. The phone is the AI computer.[LFM2.5-VL-3B — Hugging Face]( LiquidAI/LFM2.5-VL-3B)

[LFM2.5-VL-3B-GGUF — Hugging Face]( LiquidAI/LFM2.5-VL-3B-GGUF)
  • 2 replies
·