I've been working with an automatic incremental context compactor enabled and it's been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence.
TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems
I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free.
I've been working with an automatic incremental context compactor enabled and it's been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence.
TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems
404 on the blog page? https://zartbot.github.io/blog/
Yeah, for some reason their /blog/ is 404ing but for anyone interested in their other articles, https://github.com/zartbot/blog/ has all of them, just sub out anything after /blog/.* with the folder path (e.g https://zartbot.github.io/blog/arch/jalapeno/ from https://github.com/zartbot/blog/tree/main/arch/jalapeno)
Removed
The website design definitely is, but I don’t know if the content is? This reads pretty human to me and is quite interesting to boot!
you don't know him?
The writing feels human to me… and I call out AI slop as much as possible.