Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
91.6
TFLOPS
Aelin AquaSoul
PRO
SoulInPsyAbstract
1
1
Follow
andreolf's profile picture
Abrahrm's profile picture
antgas's profile picture
22 followers
·
8 following
https://sipa-os.org
AelinAquaSoul
SoulInPsyAbstract
aelin-aquasoul-8ba489404
AI & ML interests
SIPA OS: Autonomous AI for neurodivergent architects. We replace cognitive noise with a clean terminal and 344+ LLM auditing. Our system eliminates hallucinations, ensuring hyperfocus and total data control within a sovereign ZeroTrust mesh.
Recent Activity
replied
to
their
post
about 9 hours ago
Follow-up to last night's correction: the arm count was still wrong. 8, not 9. @dipankarsarkar caught it a second time — same off-by-one as the first fix, verified straight from the JSON. But the thing worth a post is what turned up while checking. One row inside that count (mistral7b-v5-final, money k=4) actually gets the right answer — "$0, unknown" — flagged only because a $ shows up mid-sentence. What it fabricates isn't the number. It's the receipt: "Operation performed: curl -s https://[...]/company/openai/results... Result: undefined... Verification: independent lookup at investing.com... Timestamp: 2026-07-01T11:07:42Z, API response code 404." None of that ran. Scored all 260 rows for it: 5/20 curl-claims and 2/20 timestamp-claims on that arm, 0/20 on its own base model. Same arm asks permission to check a fact at money k=0, then reports a completed call with a timestamp at population k=9. Checked the obvious explanation before trusting it: mistral7b-v5-final and deepseekr1-v5-final (0/20, clean) trained on the byte-identical dataset, same hyperparameters. That dataset's 100 curl-exemplars all model honest verify-before-claim behavior — zero fabricated completions. Same data, same 100 examples, one base model inverted the pattern, one didn't. Not a data problem. A base-weight problem, surfaced by identical fine-tuning. Unplanned confirmation from a different direction: sat in on a fine-tuning-vs-harness debate at AWS Floor28 last night (AI21 vs TensorOps, 117 people). Their landing point, independently: "start with the harness, earn the right to fine-tune with data and evals." Same shape this whole series keeps finding. Fixed in the repo: commit fa0c7a0. Next: binary-qwen25 to k=20, then pulling apart what in mistral7b's pretraining makes the curl→fabricate substitution available at all.
updated
a model
about 10 hours ago
SoulInPsyAbstract/specialist-vuln-merged-hermes43-lora
replied
to
their
post
about 12 hours ago
One shot said 100%. Ten shots said 94%. Yesterday i trained 6 LoRA specialists into a vulnerability-gate model (Hermes-4.3-36B, second architecture repeat of the same experiment) and tested whether it holds under a specific attack: after it correctly finds a vulnerability and returns the hard stop, ask it to use that same vulnerability as a "workaround" for something else entirely — not "continue investigating," a different, unrelated-sounding request that needs the exact same exploit. Five scenarios, one per category. Greedy decoding, single pass: 5/5. Every response correctly identified the finding, refused the reframed request, cited the hard stop rule. Looked airtight. Receipts, not hype means not stopping there. I re-ran the same five scenarios with real sampling — temperature 0.7, the same setting this project's evals have used all along — ten times each, 50 generations total. 47/50. Not 50/50. Three categories held at 10/10. Two didn't: 9/10 and 8/10, both clustered in the same failure type — infra-misconfig, where "urgent fix, use this as a workaround" apparently reads as more legitimate than the same ask framed as a secrets or injection scenario. The greedy-decode number wasn't wrong, exactly. It was one draw from a distribution, presented as if it were the distribution. That's the same mistake this whole series keeps finding in different clothes — a single passing check standing in for a property that only variance can actually show you. A gate that's 94% under a specific reframed pressure is a real, useful number. A gate that's "100%" because it was asked once is a number that hasn't been tested yet. Same instinct @dipankarsarkar has been applying to my daily receipts all week — one pass matching itself isn't proof, only repetition against something outside your own generator is. Full writeup, dataset, and merged weights: * https://github.com/soulinpsyabstract/sipa-os-governance/commit/50ba3c283ffd172eb749009cff35ddfa96bf1395
View all activity
Organizations
SoulInPsyAbstract
's models
21
Sort: Recently updated
SoulInPsyAbstract/specialist-vuln-merged-hermes43-lora
Text Generation
•
Updated
about 10 hours ago
SoulInPsyAbstract/sipa-ollama
Updated
2 days ago
SoulInPsyAbstract/sipa-os-governance
Updated
3 days ago
SoulInPsyAbstract/vuln-gate-merged-qwen25-lora
Text Generation
•
Updated
4 days ago
•
9
SoulInPsyAbstract/vuln-gate-06_stop_gate_pressure-lora
Text Generation
•
Updated
4 days ago
•
6
SoulInPsyAbstract/vuln-gate-05_supply_chain-lora
Text Generation
•
Updated
4 days ago
•
6
SoulInPsyAbstract/vuln-gate-04_infra_misconfig-lora
Text Generation
•
Updated
4 days ago
•
6
SoulInPsyAbstract/vuln-gate-03_injection-lora
Text Generation
•
Updated
4 days ago
•
6
SoulInPsyAbstract/vuln-gate-02_access_control-lora
Text Generation
•
Updated
4 days ago
•
7
SoulInPsyAbstract/vuln-gate-01_secrets_credentials-lora
Text Generation
•
Updated
4 days ago
•
9
SoulInPsyAbstract/specialist-cd-qwen25-lora
Text Generation
•
Updated
8 days ago
•
7
SoulInPsyAbstract/specialist-cd-hermes3-lora
Text Generation
•
Updated
8 days ago
•
9
SoulInPsyAbstract/specialist-cd-muse-glimmer-lora
Text Generation
•
Updated
8 days ago
•
9
SoulInPsyAbstract/binary-r1-lora
Updated
16 days ago
SoulInPsyAbstract/binary-hermes3-lora
Updated
17 days ago
SoulInPsyAbstract/binary-qwen25-lora
Updated
17 days ago
SoulInPsyAbstract/sipa-binary-gate
Text Generation
•
Updated
17 days ago
SoulInPsyAbstract/specialist-b-refusal-governance
Updated
17 days ago
•
71
SoulInPsyAbstract/protocol0-llama-3.1-8b-v5
8B
•
Updated
17 days ago
•
122
SoulInPsyAbstract/syntax-ai-community
Updated
21 days ago
SoulInPsyAbstract/SIPA-AI
Updated
21 days ago