gptel with a Custom OpenAI-Compatible Backend

In short: gptel assembles the request URL as protocol://host plus endpoint, so :host takes no scheme, :protocol silently defaults to https, and :endpoint is the whole completions path rather than a base URL. :stream defaults to false, which makes a long answer look like a hang. :key accepts a function, which is how the API key stays out of init.el, and :models is a list you write because gptel never queries /v1/models.

gptel is Emacs' general-purpose LLM client: gptel-send works in any buffer, it can rewrite a region, and in Org mode it can scope a conversation to a heading. It has no settings file and no provider dropdown in the usual sense. A backend is a Lisp object you construct, and the constructor has four defaults that decide whether your requests ever leave the machine.

Those defaults are the whole page. gptel-make-openai builds the request URL as protocol://host + endpoint — so :host must not carry a scheme, :protocol silently assumes https, and :endpoint is the entire path from the root, not a base URL. :stream defaults to nil, so a long generation arrives in one piece and looks frozen. And gptel-api-key is only one of several ways to supply the credential — the :key argument accepts a function, which is how you keep a key out of init.el entirely.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page covers where a backend is registered, what each keyword argument actually controls, why a plain-HTTP gateway fails with what looks like a dead server, and what a month of editor-side assistance costs once you know how much of it is input.

Registering the backend

Backends live in your configuration, not in a file the package reads back.

(require 'gptel)

(setq gptel-backend
      (gptel-make-openai "mygateway"      ; any name you want
        :host "gw.example.com"            ; host only — no scheme, no path
        :endpoint "/v1/chat/completions"  ; the full path from the root
        :protocol "https"                 ; this is also the default
        :stream t
        :key (lambda () (auth-source-password "gateway" "api"))
        :models '(claude-sonnet-5 gpt-5.6-sol deepseek-v4-flash)))

(setq gptel-model 'claude-sonnet-5)

With that in place, M-x gptel-send in any buffer sends the text up to point. A prefix argument (C-u C-c RET) opens the transient menu, where you can switch backend and model for the current buffer only — useful for testing a new endpoint without committing it as the default.

The URL is assembled, not given

The constructor ends with this expression:

(url (if protocol (concat protocol "://" host endpoint) (concat host endpoint)))

Everything below follows from it.

:host takes a host, not a URL

The default is "api.openai.com". Paste https://gw.example.com into :host and you get https://https://gw.example.com/v1/chat/completions. The failure looks like a DNS or TLS problem, which sends people to check their network instead of their config string. If you have a full URL, split it: scheme into :protocol, host:port into :host, path into :endpoint.

:protocol defaults to https

A gateway on http://127.0.0.1:3000 needs :protocol "http" explicitly. Leave it out and gptel builds an https:// URL against a port that speaks plain HTTP — the error text mentions TLS, so the gateway appears to be misconfigured rather than the client.

There is a documented alternative worth knowing: pass :protocol nil, and gptel uses the host string as-is, letting you put the scheme in :host yourself. Pick one convention and stay with it, because mixing the two produces exactly the doubled-scheme URL above.

:endpoint is the whole completions path

The default is "/v1/chat/completions", and it is appended verbatim. It is not a base URL: setting :endpoint "/v1" produces a request to /v1.

This is also where the version prefix belongs. In gptel's own documented backends the version lives in the endpoint, not the host — Groq's entry is :host "api.groq.com" :endpoint "/openai/v1/chat/completions", OpenRouter's is :endpoint "/api/v1/chat/completions", Mistral's is :endpoint "/v1/chat/completions". So when a vendor's snippet shows a shorter endpoint than the chat route you expect, that is a difference in the vendor's docs, not a sign that gptel appends anything. Confirm which path your gateway actually received before adjusting anything.

:stream defaults to false

(stream nil), per the constructor's own docstring: "STREAM is a boolean to toggle streaming responses, defaults to false."

For a local model this barely matters. Against a hosted reasoning model producing a long answer, non-streaming means the buffer stays empty until the last token is generated, which reads as a hang. Set :stream t unless your gateway cannot do server-sent events — and if you turn it off deliberately, note it in the config comment, because the symptom is indistinguishable from a broken endpoint.

:key accepts a function

The docstring is explicit: "KEY (optional) is a variable whose value is the API key, or function that returns the key."

That second form is the one to use. A literal string in init.el ends up in your dotfiles repository and in every backup of your home directory. A function keeps the secret where your system already stores secrets:

:key (lambda () (auth-source-password "gateway" "api"))

auth-source reads from ~/.authinfo.gpg, macOS Keychain or whatever else you have configured. The key exists only in memory at request time.

:models is your list, not the server's

gptel never calls GET /v1/models on your endpoint. The :models list you write is what fills the model menu, and gptel-model must name one of those symbols.

Two consequences that cost time:

gptel-model is a symbol, not a string, in the classic configuration. If the model menu looks right but requests fail, print the value of gptel-backend and gptel-model in the buffer you are sending from — gptel tracks a global default and a per-buffer override, and the per-buffer one wins.

Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Why the input column dominates here

An editor assistant pays for the same context repeatedly. A region rewrite re-sends the surrounding buffer; an Org-scoped conversation re-sends the heading's contents on every turn; anything you add with gptel-add-file stays in the request until you remove it. Output is one reply — usually a paragraph, a function, or a diff.

A day of work on this page's workload is 160k input and 35k output tokens — 3.2M input and 0.7M output across 20 working days. Because most gateways bill output at a multiple of input, that split is what decides the month: the input rate is the number to negotiate, and the output multiple is the number to model.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.34$6.70
gpt-5.6-sol$2.5 / $15.00$0.93$18.50
deepseek-v4-flash$0.22 / $0.66$0.06$1.17
glm-5.3$0.70 / $2.2$0.19$3.78
qwen3.8-max$1.00 / $3.00$0.27$5.30

One day = 160k input + 35k output on this page's workload. 20 days = 3.2M input and 0.7M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Verifying the route

  1. M-x gptel-send with a one-line prompt. A reply proves the key, the host, the path and the model field at once.
  2. Check gptel-backend and gptel-model in the buffer you sent from (C-h v, with the buffer current). A per-buffer override from an earlier C-u C-c RET survives longer than you expect.
  3. Read the path your gateway logged. A doubled scheme means :host carried one; a doubled /v1 means :endpoint did.
  4. Send a prompt long enough to take several seconds. Tokens appearing progressively means :stream t took effect; a single block at the end means it did not.
  5. M-x gptel-add-file a large file, then ask a question about it. That is the test that proves the editor-side context pattern works against your endpoint.

FAQ

FAQ

Why does gptel fail to connect when my gateway works in the browser?

:protocol. gptel defaults to https, and a browser hides the difference by upgrading or by warning you. Point :host at an http:// gateway without setting :protocol "http" and the TLS error names the wrong culprit.

Should `:endpoint` end in `/chat/completions`?

It should be the complete path from the host root — the default is /v1/chat/completions. gptel appends your value to protocol://host with no other logic, so anything you leave out is missing from the final URL.

Can I keep the API key out of `init.el`?

Yes — :key accepts a function, and the documented pattern is to return the value from a secret store. The alternative, gptel-api-key, is a variable whose value ends up wherever your configuration ends up.

Why is a model missing from the menu?

Because it is not in :models. gptel does not query your endpoint for a catalogue; the list you write is the list you get.

Why does a long answer appear all at once?

:stream defaults to false in gptel-make-openai. Add :stream t to the backend to receive tokens as they are produced.

Is gptel cheaper per task than an editor plugin?

The token shape is similar and the tooling is different. Both re-send buffer context, so both are input-dominated — see why output pricing dominates the bill for what the output multiple does to a month, and CodeCompanion.nvim for the Neovim equivalent of the same per-buffer problem.

Related

Get API access