MarkTechPost
AI news source. Articles are auto-selected and adapted by Hamidun News editors.
Latest publications

StepFun Releases StepAudio 2.5 Realtime Voice Model with Roleplay Support
Chinese AI lab StepFun introduced the real-time voice model StepAudio 2.5 Realtime, which outperforms competitors in speech naturalness and can adapt voice to user scenarios.

Langfuse for LLM Engineers: Complete Tracing and Experimentation Pipeline
Langfuse is a tool for debugging and optimizing LLM applications. Learn how to set up a complete monitoring pipeline, prompt management, and experiments without paid models.

WorkOS Introduces auth.md — An Open Protocol for AI Agent Registration
WorkOS has released auth.md — an open standard that enables AI agents to register in applications via Markdown file without human intervention.

ByteDance Presented Lance: One Model for Video Understanding, Generation, and Editing
ByteDance released Lance — an open-source model that operates in a single framework with images and video: understands, generates, and edits content using only 3B active parameters.

Cohere releases Command A+: 218 billion parameters for agents on two GPUs
Cohere has unveiled the open Command A+ model with 218 billion parameters and multimodal capabilities, running on two H100 GPUs and supporting 48 languages.

Perplexity Opens Bumblebee Scanner to Protect Developer Systems
Perplexity has published the source code for Bumblebee, a tool for scanning vulnerabilities in developer system dependencies without running any code.

Alibaba Introduces Qwen3.7-Max: An Agent with Million-Token Context
Alibaba introduced Qwen3.7-Max — the most advanced agent model from Qwen with a 1M-token context window and reasoning mode for complex multi-step tasks.

CopilotKit Redefines Architecture for AI Agents in 2026
CopilotKit released a new stack for agentic AI developers: AG-UI protocol, AIMock testing platform, and Pathfinder server—a complete production solution.

OpenMythos: Building Advanced Transformers with MLA and GQA in Colab
OpenMythos enables building recurrent transformers in Google Colab, comparing MLA and GQA architectures. A new video tutorial verifies model stability through spectral radius analysis of injection matrices.

Nous Research Introduces CNA: Controlling LLM Behavior Without Retraining
Nous Research introduced the Contrastive Neuron Attribution (CNA) method, which enables controlling the behavior of large language models by identifying and disabling individual neural circuits without retraining and wit

Eight best authentication platforms for AI-agents and MCP in 2026
MCP reached 97 million SDK downloads per month. AI agents are massively transitioning from experiments to production environments, and choosing the right authentication platform has become one of the…

SuperClaude Framework helps structure workflows for Claude API
SuperClaude Framework provides developers with built-in components for creating advanced AI workflows: commands, agents, execution modes, and session memory — all in one system.

Tencent Released a Local Memory System for AI Agents TencentDB
Tencent open-sourced TencentDB Agent Memory — a local memory system for AI agents that reduces token consumption by 61% and improves accuracy by 28%.

NVIDIA Introduces Gated DeltaNet-2: Linear Attention with Separate Memory Gates
NVIDIA has created a new linear attention mechanism, Gated DeltaNet-2, which improves memory management in large language models through separate erase and write gates instead of a single gate.

Google Introduces Gemini 3.5 Flash: Fast and Affordable Model for Coding and AI Agents
At I/O 2026, Google introduced Gemini 3.5 Flash — a model that is 75% cheaper than the flagship version, runs 4 times faster, and outperforms it on coding and automation tasks.

Alibaba releases a translator with 2.8-second latency across 60 languages
Alibaba introduced a model for real-time translation of video and speech simultaneously across 60 languages — with minimal latency and preservation of the speaker’s voice.

NVIDIA introduced Nemotron-Labs-Diffusion: a model with triple decoding
NVIDIA has released the Nemotron-Labs-Diffusion language model, which combines three decoding modes and processes tokens 6 times faster than Qwen3-8B.

Generating knowledge graphs from text: a practical guide with kg-gen and NetworkX
A tutorial on automatically extracting entities and relations from text with kg-gen, building interactive knowledge graphs, and analyzing them with NetworkX.

Turbovec: a Rust vector index with Google Research's TurboQuant algorithm
Turbovec uses Google's TurboQuant algorithm to compress vectors 16x without pre-training, simplifying the deployment of RAG applications.

Best platforms for agentic AI in 2026: ranking of Salesforce, Microsoft, and others
Companies are moving from pilots to production. MarkTechPost compiled a top-10 ranking of agentic AI platforms: Salesforce Agentforce, Microsoft Copilot Studio, ServiceNow, and others. Verified pricing and real-world dep

NVIDIA developed a method for training neural networks at 4-bit precision
NVIDIA introduced NVFP4, a methodology for training large models at 4-bit precision instead of the standard 8-bit, halving memory use without loss of quality.

OpenAI unveils the MRC protocol for supercomputer networks with millions of GPUs
OpenAI has created a new open network protocol, MRC, for large AI clusters. It distributes data across hundreds of paths and recovers from failures in microseconds, enabling the construction of supercomputers with 100,00

Meta AI introduces NeuralBench — a framework for testing brain activity models
Meta released NeuralBench, an open framework for standardized testing of EEG-based AI models, bringing 36 tasks, 94 datasets, and 13,603 hours of brain recordings into a single interface.

How to compress a language model 3x: a guide to FP8, GPTQ, and SmoothQuant
Developers received a step-by-step guide to compressing large language models with llmcompressor, comparing the effectiveness of FP8, GPTQ, and SmoothQuant quantization to reduce hardware load.

OpenAI released three audio models: translation, transcription, and real-time reasoning
OpenAI expanded the Realtime API with three new audio models for voice processing: reasoning agents, multilingual translation, and streaming transcription.

Anthropic created a tool to translate Claude's thoughts into human language
Anthropic developed Natural Language Autoencoders, a technology that translates Claude's internal activations into textual explanations, revealing how the neural network works.

NVIDIA packed 3 models into one file and made training 360× more efficient
NVIDIA introduced Star Elastic, a method that packs three models of different sizes into a single checkpoint and enables training that is 360× more efficient.

NVIDIA released cuda-oxide: a compiler for Rust code on GPUs
NVIDIA introduced cuda-oxide, a tool for compiling Rust functions directly into PTX code for GPUs. This will simplify the development of CUDA applications in Rust and make parallel computing more accessible.

NadirClaw: saving on LLM requests with smart prompt routing
NadirClaw is a tool for intelligent prompt routing that classifies requests as simple or complex, sending them to the appropriate model to reduce costs.

Hermes Agent by Nous Research took the lead in token usage on OpenRouter
The open-source AI agent Hermes Agent by Nous Research surpassed the closed-source platform OpenClaw and took first place on OpenRouter, generating 224 billion tokens a day. This happened in just three months and shows t