A Forecast Is More Than a Number: Evaluating Granite PatchTST-FM-r2
IBM’s latest time-series release makes a useful starting point for discussing uncertainty, honest backtests and the difference between benchmark performance and operational value.
What’s new, what matters, and what to build next.
IBM’s latest time-series release makes a useful starting point for discussing uncertainty, honest backtests and the difference between benchmark performance and operational value.
A new boundary-aware safety study highlights a familiar deployment problem: stopping harmful requests without also blocking the legitimate work an assistant exists to do.
Hcompany’s encoder release puts page images and text into the same retrieval discussion. Here is how to evaluate visual search without mistaking a retrieved page for a verified answer.
The project addressed the gap between keeping old sessions and retrieving useful context from them.
A public recipe measured whether light post-training could improve a compact model’s schema compliance.
A new JavaScript library and versioned kernel collection targeted the operations underneath browser inference.
The Monsoon evaluation sets broaden language coverage and make it possible to investigate differences hidden by an overall error rate.
Multi-vector fine-tuning offers a route to domain-specific retrieval, but good relevance data and evaluation matter as much as the model.
Typed nodes and visible intermediate outputs made multi-step AI apps easier to explore.
Speech-recognition evaluation needs to distinguish faithful listening from patterns learned around a familiar benchmark.
Explore releases and guides from 2024 onward. Dates refer to the original announcements; each article also shows when our coverage was published. How we cover the ecosystem ↗