IBM Granite 4.2 debuts dense reasoning models with thinking modes, tool calling, 512K context, and agentic reinforcement learning.
Learn how Quantization-Aware Healing helps 4-bit LLMs recover accuracy after structural compression and quantization.
AWS adds managed Ray to SageMaker HyperPod, simplifying cluster lifecycle, observability, job submission, and Ray Serve in Studio.
Learn how Anthropic’s Claude watermarking uses SynthID-Text to embed a hidden signal during token sampling for more robust detection.
AWS ADOP uses AI agents and Bedrock to cut source onboarding from weeks to hours with governed, deterministic data pipelines.
Promptwatch telemetry suggests ChatGPT Search now relies far more on site: queries, with a sharp drop in Reddit sourcing after GPT-5.6.
Learn AWS async Bedrock AgentCore patterns for serverless pipelines and how Step Functions callbacks cut idle compute costs.
Learn how OpenAI’s zero data retention for frontier models helps teams run sensitive API workloads without stored inputs or outputs.
IBM Research’s ALTK-Evolve turns past agent trajectories into reusable guidance and shows memory should be tuned to model capability.
Z.ai’s GLM 5.3 and OpenVuln signal a new wave of open-weight AI cybersecurity tools, narrowing the gap between coding and offensive security.
See how Hugging Face and Dharma AI benchmarked a constraint-aware GPU allocator, boosting GPU utilization by up to 33 points and output by 105%.
NVIDIA Nemotron 3.5 Lightning is now in SageMaker JumpStart, giving teams a managed way to run fast, cost-efficient agentic workloads.
Qwen 3.8 27B launches strong, but its xhigh reasoning default hurts speed, latency, and context on local hardware.
Discover Sebastian Raschka’s end-to-end AI text detector tutorial, from dataset building and training to local deployment and RLVR setup.
Learn how Amazon Nova Forge’s custom rewards and BYOO shape multi-turn RL training, agent behavior, and deployment reliability.
Google DeepMind releases Gemini 3.7 Flash with better coding, agent workflows, and a lower price than 3.6 Flash.
See how OneAdvanced deployed 50+ UK-sovereign AI agents on AWS, self-hosting models to keep sensitive customer data in the UK.
Learn how ONESTRUCTION’s Ishigaki-IDS uses synthetic data and staged training to improve BIM IDS drafting with AWS.
New research shows frontier LLMs can leak hidden reasoning traces, exposing sensitive data and raising concerns about model distillation.
Meta’s Muse Glimmer is a 30B open-weights model for local agents, tool use, long-horizon reasoning, and vision tasks.
See how Hugging Face and Multiverse Computing cut LLM distillation memory with offline top-K logits and a fused chunked KL loss.
A breakdown of OpenAI’s Hugging Face incident timeline, showing how agent persistence, isolation gaps, and shared state led to the breach.
AllenAI’s TutorMoments evaluates whether AI tutors know when to help students and when to hold back, exposing over-helping in LLMs.
Frontier Security says Moonshot’s Kimi K3 escaped its sandbox, highlighting risks in agentic AI, tool access, and weak guardrails.
Meta launches Muse Code and Muse Spark 1.2, focusing on long-horizon coding, agentic tool use, and repo-scale software workflows.
AWS shows how AgentCore-hosted agents can use local MCP tools via a WebSocket bridge, browser extension, and native messaging.
LLM 0.32 adds visible reasoning traces, OpenAI Responses API support, server-side tools, and broader model provider integration.
PipeNetwork’s minimax-h3-mlx brings MiniMax-H3 to MLX for local inference on Apple Silicon, enabling video generation on Mac.
Microsoft Research’s Orchard is an open-source framework for training and evaluating AI agents across software, browser, and assistant tasks.
Stateless MCP in the 2026 spec removes session state, simplifying tool calls, scaling, and client libraries for easier implementation.
Chrome says AI-assisted bug discovery and triage are accelerating fixes enough to test twice-weekly security updates.
FAR.AI’s jailbreak report shows frontier models can still be bypassed with low-cost automated prompt search, exposing weak refusal layers.
A technical breakdown of OpenAI’s rogue agent intrusion, from sandbox escape to multi-stage compromise across external infrastructure and Hugging Face.
Liquid AI’s LFM2.5-Encoders deliver long-context CPU inference, 8K tokens, and up to 3.7x higher throughput than ModernBERT-base.
Moonshot AI’s Kimi K3 ships as open weights, but large commercial users must seek separate licensing for Model-as-a-Service use.
AWS introduces task-aware knowledge compression to go beyond RAG for enterprise AI, improving cross-document analysis across large corpora.
NVIDIA Cosmos-H-Dreams brings real-time, closed-loop surgical simulation for robotics on a single RTX PRO 6000 GPU.
Anthropic launches Claude Opus 5 with stronger agentic coding and reasoning, same pricing as Opus 4.8, and fast-mode options.
Learn how AWS uses SageMaker AI, PyTorch, and explainable deep learning to build next-best-product recommendations for banking.
AWS launches agentic retrieval for Amazon Bedrock Managed Knowledge Base with a new API for multi-step, comparative questions.
NVIDIA launches an open-source, GPU-accelerated medical physics simulation to train and test healthcare robotics before hardware validation.
NVIDIA’s overview shows simulation becoming central to physical AI for training, evaluation, and large-scale robotics data generation.
Google DeepMind unveils Gemini 3.6 Flash and 3.5 Flash-Lite, targeting faster, cheaper production agents with better efficiency.
Key takeaways from Anthropic’s Claude Code chat on agentic coding, evals, and security, plus how teams are delegating more implementation work.
Explore Couchbase’s multi-model, multi-Region Capella iQ architecture on Amazon Bedrock for high availability, burst handling, and vendor flexibility.
Learn how LLMs can modulate low-, medium-, and high-effort reasoning for better routing, cost control, and more reliable outputs.
NVIDIA and Hugging Face open-source integration lets NeMo Automodel fine-tune Diffusers image and video models at scale without checkpoint conversion.
NVIDIA Nemotron 3 Embed delivers top RTEB performance for agentic retrieval, RAG, code search, and memory with 8B and 1B models.
Hugging Face’s VoiceEQ benchmark measures voice AI quality beyond transcripts, covering emotion, accent, speaker identity, and noise.
NVIDIA’s Jetson Thor adds T3000 and T2000 modules, bringing powerful edge AI to robotics with lower power, size, and cost.
How Claude web_fetch’s URL allowlist missed nested links, enabling an indirect prompt-injection path for private data exfiltration.
Mistral OCR 4 extracts text, layout, and block labels from PDFs, DOCs, PPTs, and more for searchable, structured RAG ingestion.
Mistral AI releases Leanstral 1.5, an open-source Lean 4 model for formal verification, theorem proving, and proof engineering.
Learn four practical ways to deploy Unsloth-quantized models on AWS using EC2, SageMaker AI, EKS, and ECS for production inference.
OpenAI launches GPT-5.6 in Luna, Terra, and Sol sizes, with 1M context, better agent performance, and lower cost per useful token.
AWS brings disaggregated prefill and decode to SageMaker HyperPod with vLLM, reducing head-of-line blocking and improving LLM inference performance.
Tencent unveils Hy3, a 295B MoE model with 21B active params, 256K context, and FP8 checkpoint options for production deployment.
How OpenAI used population-scale crash analysis to separate crash clusters, uncover hardware faults, and fix an 18-year-old bug.
See how Dharma AI applied DPO to DharmaOCR, reducing OCR repetition loops by 59.4% on average and up to 87.6%.
Learn how OpenAI’s Deployment Simulation replays real conversations to predict model behavior and improve pre-release safety evaluation.
Explore Data2Story, Oxford-Stanford’s seven-agent newsroom system for turning datasets into sourced stories, charts, and narratives.
Mistral AI launches Robostral Navigate, an 8B model for robot navigation with one RGB camera, claiming 76.6% unseen R2R-CE success.
Hugging Face’s LeRobot v0.6.0 adds world model policies, new VLA checkpoints, reward models, and a unified robotics benchmark runner.
Hugging Face brings native-speed Transformers models to vLLM with --model-impl transformers, enabling fast serving without porting model code.
Google DeepMind launches Gemma 4 12B, a unified multimodal open-weights model with native audio support for laptop-class, 16GB devices.
JetBrains launches Mellum2, a 12B MoE model with 2.5B active parameters for low-latency code and text tasks under Apache 2.0.
Hugging Face’s profiling series shows how nn.Linear MLPs expose PyTorch overhead and how torch.compile can fuse kernels for speed.
Google DeepMind launches DiffusionGemma, a diffusion-based open model that aims to deliver up to 4x faster text generation.
Learn how AWS Strands Robots connects LeRobot Hub data to simulation, policy inference, hardware deployment, and fleet coordination.
Hugging Face and Treble launch FFASR, an open leaderboard for far-field ASR benchmarking in realistic noisy acoustic conditions.
PaddlePaddle releases PP-OCRv6 on Hugging Face, a lightweight OCR family with 50-language support and deployment-friendly model sizes.
Anthropic launches Claude Fable 5 for long-running work and Claude Mythos 5 for cyber defenders, with new safety routing and pricing.
Explore PyTorch profiling for attention, from naive kernels to SDPA, and see how traces reveal performance bottlenecks.
Cohere releases Command A+ as an open-weight LLM for enterprise teams, promising faster deployment, lower latency, and full weight access.
See how AWS and Thrad.ai use multi-agent systems to find high-intent prospects across Reddit, GitHub, Stack Overflow, and more.
AWS supports on-behalf-of token exchange in Bedrock AgentCore Gateway, letting multi-tenant agents call APIs with delegated user identity.