Recent Posts
Testing Magnitude vs llama.cpp: Is It Really 2x Faster?
Several days ago, I came across a claim from Magnitude stating that it’s up to 2x faster than llama.cpp, and—more interestingly for me—that it supports AMD GPUs as well.
Naturally, I was skeptical. A 2x speedup over llama.cpp is a bold claim, especially on AMD hardware where the ROCm ecosystem can be hit-or-miss. So today, I decided to give it a try and run my own benchmarks.
Setup I used the following build for Magnitude:
read more
AI at scale in another sense with agent + sandbox + self-evolving + PTC
When talking about AI at scale, I have always thought about it from the model-serving infrastructure perspective:
How many GPUs do we need? How do we scale inference? How do we reduce latency and cost? How do we serve millions of concurrent requests? How do we efficiently distribute models across GPU clusters? But there is another way to think about AI at scale.
Instead of scaling the AI model itself, what if we scale software development across ordinary users?
read more
Revisiting COBOL After 26 Years
From COBOL to Modern Programming Languages About 26 years ago, I learnt COBOL using the book 《COBOL 语言》(上、下册) by 谭浩强 (Tan Haoqiang). Here is a portion of my transcript.
At the time, I developed a simple MIS (Management Information System) using COBOL. It was one of the programming languages I worked with early in my career.
Since then, I never really had the opportunity to work with COBOL professionally.
read more