NEAR AI Cloud is now available as an inference provider on OpenRouter, giving developers access to its model-serving infrastructure through an existing OpenRouter API key and balance. According to NEAR AI’s official September 28 announcement, the service appears under the provider slug near-ai and initially serves GLM 5.3 Flash. The integration expands distribution for NEAR AI Cloud, but it does not extend the platform’s confidential-inference guarantee through OpenRouter.
GLM 5.3 Flash is a multimodal model from Z.ai designed for coding, agentic workflows, reasoning, tool use and visual understanding. NEAR AI lists a 1,048K-token context window for the model, alongside pricing of $0.15 per million input tokens and $0.50 per million output tokens on its direct service. OpenRouter users can now reach the same NEAR AI-hosted model without opening a separate NEAR AI account, reducing integration friction for developers already using the router.
OpenRouter Adds Distribution, Not Confidential Inference
The privacy boundary is the most important part of the integration. NEAR AI says GLM 5.3 Flash still runs on the same Trusted Execution Environment hardware used by its direct API, but requests arriving through OpenRouter first traverse infrastructure outside that confidential environment. The OpenRouter gateway path is not attested, so users cannot treat an OpenRouter request as end-to-end confidential inference simply because NEAR AI ultimately serves the model.
That distinction builds on NEAR AI’s broader confidential inference architecture with hardware-backed attestations. Direct NEAR AI requests to supported open-weight models run inside Intel TDX virtual machines paired with NVIDIA GPUs operating in confidential-computing mode. Attestation reports allow users to verify the execution environment independently. Confidential computing protects workloads while they are being processed, but only when the complete relevant request path remains inside the intended trust boundary.
OpenRouter introduces another routing consideration: provider fallback. NEAR AI can be prioritized by placing near-ai first in OpenRouter’s provider order, but fallbacks remain enabled by default. If NEAR AI cannot serve the request, OpenRouter may route it to another compatible provider unless the user explicitly disables fallback behavior. Developers requiring a specific execution provider must therefore pin NEAR AI and turn off fallbacks rather than assuming every request reaches the same infrastructure.
NEAR AI Builds Multiple Access Paths
The integration follows a broader expansion of NEAR AI’s OpenAI-compatible model-access layer, which lets developers use familiar client libraries while selecting among different privacy and execution models. NEAR AI separately offers confidential open-weight inference, externally routed frontier models and integrations with privacy-oriented gateways. The common API surface simplifies model access, but privacy remains a property of the selected route rather than a universal feature attached to every request.
That separation is particularly relevant as NEAR assembles cross-chain execution, private inference and AI-agent infrastructure into a broader agent stack. Autonomous systems may need model inference, payments and external tools through multiple providers, making execution provenance increasingly important. An agent can gain broader model availability through a router while still needing a direct, attested channel for workloads containing sensitive data or proprietary instructions.
The next meaningful milestone will be measurable OpenRouter usage through the near-ai provider and expansion beyond the single GLM 5.3 Flash listing. For now, the integration is best understood as a distribution upgrade: OpenRouter broadens access to NEAR AI Cloud, while confidential inference continues to require a route whose entire relevant gateway and compute path can be verified.