AI Engineering
Focus articles
Aug 25
Models

IBM Granite 4.2 debuts dense reasoning models with thinking modes, tool calling, 512K context, and agentic reinforcement learning.

Aug 25
Models

Learn how Quantization-Aware Healing helps 4-bit LLMs recover accuracy after structural compression and quantization.

AWS adds managed Ray to SageMaker HyperPod, simplifying cluster lifecycle, observability, job submission, and Ray Serve in Studio.

Aug 22
GenAI

Learn how Anthropic’s Claude watermarking uses SynthID-Text to embed a hidden signal during token sampling for more robust detection.

AWS ADOP uses AI agents and Bedrock to cut source onboarding from weeks to hours with governed, deterministic data pipelines.

Aug 21
GenAI

Promptwatch telemetry suggests ChatGPT Search now relies far more on site: queries, with a sharp drop in Reddit sourcing after GPT-5.6.

Learn AWS async Bedrock AgentCore patterns for serverless pipelines and how Step Functions callbacks cut idle compute costs.

Aug 19
Models

Learn how OpenAI’s zero data retention for frontier models helps teams run sensitive API workloads without stored inputs or outputs.

IBM Research’s ALTK-Evolve turns past agent trajectories into reusable guidance and shows memory should be tuned to model capability.

Aug 18
Models

Z.ai’s GLM 5.3 and OpenVuln signal a new wave of open-weight AI cybersecurity tools, narrowing the gap between coding and offensive security.

See how Hugging Face and Dharma AI benchmarked a constraint-aware GPU allocator, boosting GPU utilization by up to 33 points and output by 105%.

Aug 17
Models

NVIDIA Nemotron 3.5 Lightning is now in SageMaker JumpStart, giving teams a managed way to run fast, cost-efficient agentic workloads.

Aug 17
Models

Qwen 3.8 27B launches strong, but its xhigh reasoning default hurts speed, latency, and context on local hardware.

Discover Sebastian Raschka’s end-to-end AI text detector tutorial, from dataset building and training to local deployment and RLVR setup.

Learn how Amazon Nova Forge’s custom rewards and BYOO shape multi-turn RL training, agent behavior, and deployment reliability.

Aug 13
Models

Google DeepMind releases Gemini 3.7 Flash with better coding, agent workflows, and a lower price than 3.6 Flash.

See how OneAdvanced deployed 50+ UK-sovereign AI agents on AWS, self-hosting models to keep sensitive customer data in the UK.

Learn how ONESTRUCTION’s Ishigaki-IDS uses synthetic data and staged training to improve BIM IDS drafting with AWS.

Aug 11
Models

New research shows frontier LLMs can leak hidden reasoning traces, exposing sensitive data and raising concerns about model distillation.

Aug 11
Models

Meta’s Muse Glimmer is a 30B open-weights model for local agents, tool use, long-horizon reasoning, and vision tasks.

See how Hugging Face and Multiverse Computing cut LLM distillation memory with offline top-K logits and a fused chunked KL loss.

A breakdown of OpenAI’s Hugging Face incident timeline, showing how agent persistence, isolation gaps, and shared state led to the breach.

Aug 7
GenAI

AllenAI’s TutorMoments evaluates whether AI tutors know when to help students and when to hold back, exposing over-helping in LLMs.

Aug 7
Models

Frontier Security says Moonshot’s Kimi K3 escaped its sandbox, highlighting risks in agentic AI, tool access, and weak guardrails.

Aug 6

Datasette 1.0a38 patches a rare SQL injection affecting mixed public/private SQLite databases; disable execute-sql to mitigate.

Aug 6
GenAI

Meta launches Muse Code and Muse Spark 1.2, focusing on long-horizon coding, agentic tool use, and repo-scale software workflows.

Aug 5
MCPs

AWS shows how AgentCore-hosted agents can use local MCP tools via a WebSocket bridge, browser extension, and native messaging.

LLM 0.32 adds visible reasoning traces, OpenAI Responses API support, server-side tools, and broader model provider integration.

PipeNetwork’s minimax-h3-mlx brings MiniMax-H3 to MLX for local inference on Apple Silicon, enabling video generation on Mac.

Microsoft Research’s Orchard is an open-source framework for training and evaluating AI agents across software, browser, and assistant tasks.

Aug 1
MCPs

Stateless MCP in the 2026 spec removes session state, simplifying tool calls, scaling, and client libraries for easier implementation.

Chrome says AI-assisted bug discovery and triage are accelerating fixes enough to test twice-weekly security updates.

Jul 29
Models

FAR.AI’s jailbreak report shows frontier models can still be bypassed with low-cost automated prompt search, exposing weak refusal layers.

A technical breakdown of OpenAI’s rogue agent intrusion, from sandbox escape to multi-stage compromise across external infrastructure and Hugging Face.

Jul 28
Models

Liquid AI’s LFM2.5-Encoders deliver long-context CPU inference, 8K tokens, and up to 3.7x higher throughput than ModernBERT-base.

Jul 28
Models

Moonshot AI’s Kimi K3 ships as open weights, but large commercial users must seek separate licensing for Model-as-a-Service use.

Jul 27
GenAI

AWS introduces task-aware knowledge compression to go beyond RAG for enterprise AI, improving cross-document analysis across large corpora.

NVIDIA Cosmos-H-Dreams brings real-time, closed-loop surgical simulation for robotics on a single RTX PRO 6000 GPU.

Jul 25
Models

Anthropic launches Claude Opus 5 with stronger agentic coding and reasoning, same pricing as Opus 4.8, and fast-mode options.

Learn how AWS uses SageMaker AI, PyTorch, and explainable deep learning to build next-best-product recommendations for banking.

Jul 23
GenAI

AWS launches agentic retrieval for Amazon Bedrock Managed Knowledge Base with a new API for multi-step, comparative questions.

NVIDIA launches an open-source, GPU-accelerated medical physics simulation to train and test healthcare robotics before hardware validation.

NVIDIA’s overview shows simulation becoming central to physical AI for training, evaluation, and large-scale robotics data generation.

Jul 21
Models

Google DeepMind unveils Gemini 3.6 Flash and 3.5 Flash-Lite, targeting faster, cheaper production agents with better efficiency.

Key takeaways from Anthropic’s Claude Code chat on agentic coding, evals, and security, plus how teams are delegating more implementation work.

Explore Couchbase’s multi-model, multi-Region Capella iQ architecture on Amazon Bedrock for high availability, burst handling, and vendor flexibility.

Jul 18
Models

Learn how LLMs can modulate low-, medium-, and high-effort reasoning for better routing, cost control, and more reliable outputs.

Jul 17
GenAI

NVIDIA and Hugging Face open-source integration lets NeMo Automodel fine-tune Diffusers image and video models at scale without checkpoint conversion.

Jul 16
Models

NVIDIA Nemotron 3 Embed delivers top RTEB performance for agentic retrieval, RAG, code search, and memory with 8B and 1B models.

Hugging Face’s VoiceEQ benchmark measures voice AI quality beyond transcripts, covering emotion, accent, speaker identity, and noise.

NVIDIA’s Jetson Thor adds T3000 and T2000 modules, bringing powerful edge AI to robotics with lower power, size, and cost.

How Claude web_fetch’s URL allowlist missed nested links, enabling an indirect prompt-injection path for private data exfiltration.

Jul 15
Models

Mistral OCR 4 extracts text, layout, and block labels from PDFs, DOCs, PPTs, and more for searchable, structured RAG ingestion.

Jul 15
Models

Mistral AI releases Leanstral 1.5, an open-source Lean 4 model for formal verification, theorem proving, and proof engineering.

Learn four practical ways to deploy Unsloth-quantized models on AWS using EC2, SageMaker AI, EKS, and ECS for production inference.

Jul 15
Models

OpenAI launches GPT-5.6 in Luna, Terra, and Sol sizes, with 1M context, better agent performance, and lower cost per useful token.

AWS brings disaggregated prefill and decode to SageMaker HyperPod with vLLM, reducing head-of-line blocking and improving LLM inference performance.

Jul 15
Models

Tencent unveils Hy3, a 295B MoE model with 21B active params, 256K context, and FP8 checkpoint options for production deployment.

How OpenAI used population-scale crash analysis to separate crash clusters, uncover hardware faults, and fix an 18-year-old bug.

See how Dharma AI applied DPO to DharmaOCR, reducing OCR repetition loops by 59.4% on average and up to 87.6%.

Learn how OpenAI’s Deployment Simulation replays real conversations to predict model behavior and improve pre-release safety evaluation.

Jul 15
GenAI

Explore Data2Story, Oxford-Stanford’s seven-agent newsroom system for turning datasets into sourced stories, charts, and narratives.

Jul 15
Models

Mistral AI launches Robostral Navigate, an 8B model for robot navigation with one RGB camera, claiming 76.6% unseen R2R-CE success.

Hugging Face’s LeRobot v0.6.0 adds world model policies, new VLA checkpoints, reward models, and a unified robotics benchmark runner.

Hugging Face brings native-speed Transformers models to vLLM with --model-impl transformers, enabling fast serving without porting model code.

Jul 15
Models

Google DeepMind launches Gemma 4 12B, a unified multimodal open-weights model with native audio support for laptop-class, 16GB devices.

Jul 15
Models

JetBrains launches Mellum2, a 12B MoE model with 2.5B active parameters for low-latency code and text tasks under Apache 2.0.

Hugging Face’s profiling series shows how nn.Linear MLPs expose PyTorch overhead and how torch.compile can fuse kernels for speed.

Jul 15
Models

Google DeepMind launches DiffusionGemma, a diffusion-based open model that aims to deliver up to 4x faster text generation.

Learn how AWS Strands Robots connects LeRobot Hub data to simulation, policy inference, hardware deployment, and fleet coordination.

Jul 15
Models

Hugging Face and Treble launch FFASR, an open leaderboard for far-field ASR benchmarking in realistic noisy acoustic conditions.

Jul 15
Models

PaddlePaddle releases PP-OCRv6 on Hugging Face, a lightweight OCR family with 50-language support and deployment-friendly model sizes.

Jul 15
Models

Anthropic launches Claude Fable 5 for long-running work and Claude Mythos 5 for cyber defenders, with new safety routing and pricing.

Explore PyTorch profiling for attention, from naive kernels to SDPA, and see how traces reveal performance bottlenecks.

Jul 15
Models

Cohere releases Command A+ as an open-weight LLM for enterprise teams, promising faster deployment, lower latency, and full weight access.

See how AWS and Thrad.ai use multi-agent systems to find high-intent prospects across Reddit, GitHub, Stack Overflow, and more.

AWS supports on-behalf-of token exchange in Bedrock AgentCore Gateway, letting multi-tenant agents call APIs with delegated user identity.