Recent Posts
DSH: first try
deepseek harness is quite promising considering MIT license, plugin-based architecutre and harness optimization. Today I tried to work on how to write plugins for cordis and deepseek-harness. In the past I built it from source and run it. however I got below issues this time when I tried to develop somethings following develop guide.
# following the steps mentioned in readme.md pnpm install pnpm run build ERROR Error: [@deepseek-ai/dsh-root] Cannot find entry: ["lib/types/{index,invariant,startup}.
read more
Testing Magnitude vs llama.cpp: Is It Really 2x Faster?
Several days ago, I came across a claim from Magnitude stating that it’s up to 2x faster than llama.cpp, and—more interestingly for me—that it supports AMD GPUs as well.
Naturally, I was skeptical. A 2x speedup over llama.cpp is a bold claim, especially on AMD hardware where the ROCm ecosystem can be hit-or-miss. So today, I decided to give it a try and run my own benchmarks.
Setup I used the following build for Magnitude:
read more
AI at scale in another sense with agent + sandbox + self-evolving + PTC
When talking about AI at scale, I have always thought about it from the model-serving infrastructure perspective:
How many GPUs do we need? How do we scale inference? How do we reduce latency and cost? How do we serve millions of concurrent requests? How do we efficiently distribute models across GPU clusters? But there is another way to think about AI at scale.
Instead of scaling the AI model itself, what if we scale software development across ordinary users?
read more