Semantic Kernel with a Custom OpenAI-Compatible Endpoint

In short: The answer differs by language. .NET's AddOpenAIChatCompletion takes an endpoint argument, but custom endpoints are flagged experimental — the build fails until SKEXP0010 is suppressed. Python's OpenAIChatCompletion has no base URL parameter at all; the route in is an AsyncOpenAI client passed as async_client.

Semantic Kernel is Microsoft's SDK for wiring models into an application — .NET, Python and Java, with plugins, planners and dependency injection around the call. Pointing it at your own endpoint is normally a matter of reusing its OpenAI connector, because the OpenAI chat completions shape is what most gateways speak.

The catch is that the answer is different per language, and any single answer is wrong for one of them. .NET's AddOpenAIChatCompletion takes an endpoint argument — but Microsoft's own documentation marks custom endpoints on the OpenAI connector as experimental, so the build fails until you suppress SKEXP0010. Python's OpenAIChatCompletion has no base URL parameter at all; the way in is to construct an AsyncOpenAI client yourself and hand it over as async_client. And in both, the model id you pass is the string that gets sent, which makes it the first thing to check when a request returns a model error.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page covers both routes with working code, why the experimental pragma exists and what happens without it, why the Python path needs an async client specifically, how to pick a service when you register more than one, and what a month of server-side traffic costs when your prompts carry retrieved context.

.NET: endpoint is an argument, and it is experimental

using Microsoft.SemanticKernel;

#pragma warning disable SKEXP0010

IKernelBuilder builder = Kernel.CreateBuilder();
builder.AddOpenAIChatCompletion(
    modelId: "claude-sonnet-5",
    apiKey: Environment.GetEnvironmentVariable("GATEWAY_API_KEY"),
    endpoint: new Uri("https://your-gateway.example.com/v1"),
    serviceId: "gateway"
);
Kernel kernel = builder.Build();

Two things about this that are not obvious:

The pragma is not optional. Microsoft documents custom endpoints on the OpenAI connector as experimental, and the guidance is to add #pragma warning disable SKEXP0010 to use it. In a project with TreatWarningsAsErrors, the build simply fails — which reads like a broken package reference rather than an intentional flag.

The argument is a Uri, not a string. A string will not bind, and the overload you want is the one with all four of modelId, apiKey, endpoint and optionally serviceId. The same shape exists on builder.Services for dependency injection, and as new OpenAIChatCompletionService(..., endpoint: new Uri(...)) if you want to construct the service directly.

Python: there is no base URL parameter

OpenAIChatCompletion accepts ai_model_id, service_id, api_key, org_id, default_headers, async_client, env_file_path, env_file_encoding and instruction_role. There is no base_url, no endpoint, no api_base. The way in is async_client:

import os
from openai import AsyncOpenAI
from semantic_kernel import Kernel
from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion

kernel = Kernel()
kernel.add_service(
    OpenAIChatCompletion(
        ai_model_id="claude-sonnet-5",
        service_id="gateway",
        async_client=AsyncOpenAI(
            base_url="https://your-gateway.example.com/v1",
            api_key=os.environ["GATEWAY_API_KEY"],
        ),
    )
)

Two consequences:

It must be AsyncOpenAI. The parameter is typed for the async client; a synchronous OpenAI instance is not a drop-in here, and the failure surfaces later, when the kernel tries to await a call.

The key can be set in two places and they can disagree. You can pass api_key to the Semantic Kernel service and to the client. The client is the thing that actually signs the request, so put the credential on AsyncOpenAI and treat the service-level argument as decoration you do not need.

The model id is the string that gets sent

modelId in .NET and ai_model_id in Python are passed through to the request body as the model field. They are not deployment names, display names or aliases unless your gateway happens to accept one.

This is the single most common cause of a request that authenticates fine and then returns a model error: the endpoint accepted your key, which proves routing, and rejected the model, which proves nothing about the URL. Copy the id from your gateway's model list rather than from a vendor's marketing page — model ids in this catalogue frequently carry a slash, and the slash is part of the id, not a path.

Registering more than one service

Once two chat services exist, "give me the chat service" is ambiguous:

var chat = kernel.GetRequiredService<IChatCompletionService>("gateway");
chat = kernel.get_service(service_id="gateway")

Without a service_id, which one you get is not something to leave to the container — particularly if one service points at a gateway and another at a vendor directly, and your cost assumptions depend on which ran.

Both routes land on POST /v1/chat/completions

Whether you configure .NET or Python, the request is the same, which gives you one test that covers both:

# gateway log, most recent request
POST /v1/chat/completions           ← correct
POST /v1/v1/chat/completions        ← your base URL already ends in /v1 and the client adds it
POST /v1/embeddings                 ← wrong service wired up
Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Why the input column dominates here

Semantic Kernel applications are usually server-side, and the pattern that shows up over and over is retrieval: you assemble a prompt from documents you looked up, then ask for a comparatively short answer. The retrieved chunk is large and the reply is small, so the input column carries the month.

A day of work on this page's workload is 300k input and 40k output tokens — 6M input and 0.8M output across 20 working days. That is the most input-heavy split on this site, which means two things in practice: the input rate is what to optimise for, and prompt caching is worth more here than on any chat-shaped workload, because the same retrieved context is often re-sent across requests.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.50$10.00
gpt-5.6-sol$2.5 / $15.00$1.35$27.00
deepseek-v4-flash$0.22 / $0.66$0.09$1.85
glm-5.3$0.70 / $2.2$0.30$5.96
kimi-k3$1.5 / $7.5$0.75$15.00

One day = 300k input + 40k output on this page's workload. 20 days = 6M input and 0.8M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Verifying the route

  1. Send one request and read the gateway log. Confirm the path is /v1/chat/completions and the status is 200 — that proves URL, key and model id together.
  2. If it is a 404 with a doubled /v1, drop the suffix from your base URL.
  3. If it is a 401, check the credential reached the client, not just the service.
  4. If it is a 200 with an empty body, you are probably hitting the wrong service — an embeddings route, or a second registered service without the service_id you asked for.

FAQ

FAQ

Why does my build fail with SKEXP0010?

Custom endpoints on Semantic Kernel's OpenAI connector are documented as experimental. Add #pragma warning disable SKEXP0010 above the registration, or disable that specific warning in the project file.

Why is there no `base_url` argument on the Python service?

There is not one. Construct AsyncOpenAI(base_url=..., api_key=...) and pass it as async_client.

Do I need the async client specifically?

Yes — the parameter expects an AsyncOpenAI. A synchronous client does not satisfy it.

Which model id should I pass?

The exact id your gateway serves. It is sent verbatim as the model field, including any slash. A 200 on auth followed by a model error means the URL is right and the id is not.

How do I keep two endpoints apart?

Give each a service_id and request it explicitly. Without one, resolution is not something to rely on when cost depends on which service answered.

Does this work with function calling and plugins?

The request is standard OpenAI chat completions, so tool schemas pass through. What varies is whether the model you named supports them — a capability question, not a routing one.

Related

Get API access