Linked Zero Sync 
ai & security; research and policy @ big corp. pnw 🏔️🌲
There are two readings here:
- optimising efficiency means reduced costs (fiscal & environmental)
- optimising efficiency means increased consumption (jevon's paradox)
And a very cynical third one:
- Money not spent on infra is money that can be billed for compute facilities
I think both are valid. But either way, doing the same for less is for some reason attractive to me.
I don't know how many ways to say this but
- stop looking at model performance
- it's almost all in the harness
- decent harness can easily outclass top model
evidence item #71625: https://developer.nvidia.com/blog/create-a-langchain-deep-agents-harness-profile-for-nvidia-nemotron-3-ultra-to-improve-performance/
I want to jump on a couple of false dichotomies around LLM speak:
- "frontier" models vs. open model - leading models can be open
- closed model vs. Chinese model - where the model's made has no impact on how it's distributed
You can have open frontier models, closed Chinese models, open US models, closed non-frontier models (private models make sense!)
-
Falsification Engine: After finding a potential vulnerability, VulnHunter runs a structured reasoning workflow specifically designed to disprove its own argument. It searches for flawed assumptions, logic gaps, or security controls that would block the attack.
-
Evidence-Backed Remediation: When a defect survives the falsification engine, VulnHunter maps the exact exploit path and generates focused, targeted code changes for review.
-
Attacker-First Forward Analysis: VulnHunter flips the "sink-first" security model to reduce false positives by simulating a bad actor's exact journey. It begins at potential attacker-accessible entry points (APIs, network messages, file uploads) and reasons forward to evaluate whether an attacker can truly break through.
Solid US-origin open model, congratulations Thinking Machines
- context window of 1M
- available on hugging face now
- between opus 4.6 and gpt 5.6 on a web dev benchmark
- token efficient
- 41B active params of 975B total
https://www.wired.com/story/thinking-machines-lab-releases-its-first-model-inkling/
Anonymised analysis of the openai model 'breaching' hugging face:
report doesn't say what sandbox sol broke out of??
a docker container running as root
Plot twist there was no sandbox at all
many use "sandbox" and "container with host access" interchangeably
ymmv, use critical thinking

