RepoSkein serves a deterministic code graph to your agent over MCP. No database, no daemon, versioned with git, claims 8.4x fewer context tokens than grep-and-guess. Static analysis beats LLM extraction when you need reproducible.
Jeremy Morgan
I'm Jeremy, and I help developers get better at what they do.
Geek, Developer, Tech Blogger, and Volunteer Firefighter.
I once held the world record for being the youngest person alive.
Super happy with my latest Linux Foundation (@linuxfoundation@social.lfx.dev) certification. I would definitely recommend it, it covers some important topics. AMA
LiteLLM versions 1.82.7 and 1.82.8 on PyPI were compromised. The payload executed at Python startup with no import required. If your CI ever ran unpinned pip install litellm, treat it like full host compromise.
Ghostty 1.3 fixed a memory leak where Claude Code's heavy output could consume up to 37GB over ten days. If you run AI coding agents that produce heavy terminal output, this fix alone justifies upgrading.
RAGFlow added checkpoint/resume to community extraction and entity resolution, the most expensive parts of GraphRAG indexing. When your problems become resumability and state repair, the tech is maturing. "Start over" was never a recovery strategy.
Text-to-SQL agents fail before SQL generation: they don't know which table means what or which joins matter. Neocarta builds a graph from BigQuery metadata, business terms, and observed query logs, then serves it to agents over MCP. Attacks the failure upstream.
Most hallucination advice is vibes. This is five techniques with before/after demos: GraphRAG for verifiable aggregation, symbolic guardrails, runtime self-correction. Tool selection alone drops token usage from thousands to under 300.
Anthropic chose not to release its latest model after it found thousands of vulnerabilities in popular software, some dating back 27 years. The US Treasury gathered banking leaders in Washington to discuss the cybersecurity fallout. Whether this is real or safety-washing, the risks to old codebases are now urgent.
axios Compromised on npm - Malicious Versions Drop Remote Access Trojan
If you're building stuff with Langchain, you need Langsmith tools. Even if you use it for nothing else, monitoring costs in near realtime is worth the few minutes to set it up.
9 RAG Architectures Every AI Developer Must Know: A Complete Guide with Examples
A Harness report found that AI coding tools increase deployment frequency while incident recovery time gets worse and failures rise. More PRs is not the same thing as more reliability. The bottleneck moves to CI, review, and rollback automation.
A new paper removes the CPU from LLM inference entirely. Blink uses a SmartNIC to deliver inputs directly into GPU memory via RDMA while a persistent GPU kernel handles batching and scheduling. P99 time-to-first-token drops by up to 8.47x. Under CPU interference, existing systems degrade by two orders of magnitude. Blink stays stable.
A GraphRAG ablation where the graph doesn't automatically win. Mandatory KG access: 94%. No KG: 100% but fabricated material properties. Adaptive retrieval: 100% with near-perfect fidelity. The lesson is retrieval policy, not better retrieval.
FalkorDB patch worth reading: bulk deletions were failing silently. If you run it in production, upgrade this week and check whether your retention jobs actually deleted anything. Silent write failures are the worst database failure class.
One microservice catalog, three knowledge bases: Markdown, GraphRAG, exact graph tools over MCP. Markdown missed consumers, GraphRAG pulled in unaffected neighbors, graph tools computed dependencies deterministically. The failure shapes are the finding.
Full implementation of PDE-Agents: LangGraph, local LLMs, Neo4j GraphRAG, FEniCSx, plus the complete 50-task eval stack. Rare to get an agent paper you can actually rerun.
You already know more Cypher than you think.
I translated 10 familiar SQL queries into Cypher, from SELECT and JOIN to recursive dependency traversal, and showed where graph queries become simpler.
https://graphacademy.neo4j.com/blog/sql-to-cypher-10-queries
Meta published a rare public admission that they let jemalloc drift from core engineering principles and accumulate serious technical debt. jemalloc is the default allocator for Redis, FreeBSD, Firefox, and hundreds of other projects. Their problem is your problem.
CodeGraph 1.0 fixed a symlink traversal exposing files outside the indexed root and stopped surfacing Spring config secrets to agents. Once an agent traverses your code graph, edge correctness is a security property. Read release notes accordingly.
The repo behind the Markdown vs GraphRAG vs graph-tools experiment: both services, all three KBs, the eval tasks, the AI judge. Rerun it on your own architecture instead of trusting a two-service result.
LLM-generated knowledge graphs are cheap. Trustworthy ones aren't. myKG puts confidence and provenance on every edge and leaves uncertain values unknown instead of promoting guesses to facts. The boring trust machinery every KG demo postpones.
Text-to-SQL agents fail before SQL generation: they don't know which table means what or which joins matter. Neocarta builds a graph from BigQuery metadata, business terms, and observed query logs, then serves it to agents over MCP. Attacks the failure upstream.
Build a voice assistant that lets you ask natural-language questions to a Neo4j graph database and hear the answers spoken back.
eBPF.party lets you write and run eBPF programs from a browser tab. Each run gets a disposable Firecracker VM with no network, 64MB RAM, and a 500ms self-destruct timer. The setup tax for learning eBPF just dropped to zero.
GitHub Copilot CLI now lets you run two different foundation models on the same task. Claude Sonnet generates the plan, GPT-5.4 reviews it before execution. The combo closed about 75% of the performance gap on complex multi-file changes. Cross-model verification is becoming the new default.
Anthropic shelved a model after it found thousands of legacy zero-days. The US Treasury called bank execs to Washington. Also this week: GitHub Copilot CLI now cross-checks with a second model before executing, and Infercache makes long-context inference actually affordable. New issue is live.
Fully offline agentic GraphRAG for your codebase. Ask "what breaks if I change this function" and get file paths and line numbers, zero API calls. Codebase Q&A is where multi-hop retrieval stops being theoretical.
Vector Databases Explained: The Infrastructure Behind Modern AI
I am blessed. I have a fast machine with an RTX4090 and a bunch of RAM. During the day it runs Linux and I use it to work mostly thru SSH. I've build so much with it.
Some nights, I reboot it into Windows and it becomes a incredible gaming machine for competitive sim racing!
This poor thing never gets a break.
The Junior Developer Pipeline Is Broken... And AI Broke It
https://dev.to/nazar-boyko/the-junior-developer-pipeline-is-broken-and-ai-broke-it-1aai
Researchers audited a large set of AI agent "skills" and found a meaningful fraction containing hidden malicious directives like data exfiltration. Natural-language instruction packs are the new untyped, unreviewed dependency. Treat them like npm packages.
Cloudflare traced 30-minute Atlantis restarts to one Kubernetes field. kubelet was recursively changing group ownership on a persistent volume with millions of files. One-line fix, restart dropped to 30 seconds, 600 engineering hours saved per year.
https://blog.cloudflare.com/one-line-kubernetes-fix-saved-600-hours-a-year/
Han Xiao open-sourced a KG extractor running Qwen on a single L4. Docs in, streaming triples out, evidence spans and confidence per edge. The prompting tricks that force canonical entities are a free lesson for anyone building extraction without API bills.
The Schema Is the Product: An Architectural Reading of Karpathy’s LLM Wiki
A cuBLAS bug is causing RTX 5090 GPUs to use only about 40% of available compute on batched FP32 matrix multiply workloads. If you bought one for local AI work, you are getting less than half the performance you paid for. NVIDIA has not acknowledged it publicly yet.
Karpathy sketched the gap. Graphify turned it into a CLI. A persistent knowledge-graph layer for coding agents that claims up to 71.5x fewer query tokens on mixed corpora. 23.5k GitHub stars already. The "read the graph before grep" pattern is the kind of small workflow upgrade that makes coding assistants less wasteful on large repos.
This is the best lightning video from the Portland area storm this weekend. My wife filmed this from the couch! This was going on in our front yard.
https://www.facebook.com/share/r/18TZd328AF/?mibextid=wwXIfr
Someone ran a 397 billion parameter model on a MacBook with 48GB of RAM. No Python. No frameworks. Just C and Metal shaders. The trick: MoE models activate only a few experts per token, so you can stream a 209GB model from SSD with only 5.5GB in memory.
The fastest structured path to a graph-backed agent: a few focused hours, free, on Aura. This is the exact skill every GraphRAG paper this quarter keeps circling. Worth doing before your next architecture review.
Your eight-GPU box might be CPU-bound. A new paper analyzed 4.65 million scheduler records and found H100 jobs running with as little as 1 CPU core for 8 GPUs. Adding enough CPU lifted time-to-first-token by up to 5.4x without adding any GPUs.
Your GraphRAG isn't failing at retrieval. It's failing at construction. KDD 2026 paper shows graph RAG boosts recall while cratering relevance, then fixes it with conflict-resolving agents. Deleting 40% of low-frequency triples improved accuracy. Audit your KG.
