How to Get a YouTube Transcript with a REST API (curl, Webhooks, HMAC)
Short answer: send a POST with the YouTube link, wait until the job is done (poll it, or let a webhook tell you), then GET the transcript as json, text, srt or vtt. That is three calls. This post shows each call with curl, then adds webhooks with a signature check in Node.js and Python.
I build Libraryminds, so the examples use its v1 API. Everything below comes from the public developer docs, and the code blocks are copied from there unless I say I changed them. A new account gets 190 free minutes with no card, which is enough to try every call here.
What you need first
An account and an API key. Keys start with lm_live_ and you create them in account settings.
Every request sends the key as a Bearer token:
Authorization: Bearer lm_live_YOUR_API_KEYThe base URL is https://libraryminds.com and it is HTTPS only.
The default limit is 60 requests per minute per API key. AI endpoints allow 30 requests per 15 minutes, and chat endpoints allow 20 per 5 minutes.
Keys have scopes. A key without the scope for a call gets 403 insufficient_scope.
The scopes are account:read, jobs:read, jobs:write, transcripts:read, ai:write, search:read, exports:read, usage:read, webhooks:read and webhooks:write. You can have up to 10 keys per account, so make one key per app and give it only the scopes that app needs.
Step 1: Submit the video
curl -X POST https://libraryminds.com/api/v1/jobs \
-H "Authorization: Bearer lm_live_YOUR_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: my-unique-key-123" \
-d '{"url": "https://youtube.com/watch?v=dQw4w9WgXcQ"}'
Jobs are asynchronous. The API answers 202 Accepted, and the transcript is not ready yet.
The Idempotency-Key header is optional. If you send the same key with the same body again, you get the original response back. If you reuse the key with a different body, you get 422. The key can be up to 255 characters. This is what makes a retry after a network timeout safe to send.
Two other ways to start a job:
curl -X POST https://libraryminds.com/api/v1/jobs/upload \
-H "Authorization: Bearer lm_live_YOUR_API_KEY" \
-F "file=@meeting.mp3" \
-F "language=en"
POST /api/v1/jobs/import-url takes a direct HTTPS link to a media file. Languages are auto-detected, and the docs list 10: English, Hindi, Spanish, French, German, Portuguese, Mandarin Chinese, Japanese, Korean and Arabic.
Step 2: Check the job
curl https://libraryminds.com/api/v1/jobs/JOB_ID \
-H "Authorization: Bearer lm_live_YOUR_API_KEY"
JOB_ID is the id of the job you just created. You can also list jobs with GET /api/v1/jobs (it can filter by status), cancel a pending or processing job with POST /api/v1/jobs/:id/cancel, and retry a failed one with POST /api/v1/jobs/:id/retry.
If you poll, keep the 60 requests per minute limit in mind. A check every 5 to 10 seconds is plenty. Or skip polling and use a webhook, shown below.
Step 3: Get the transcript
curl "https://libraryminds.com/api/v1/jobs/JOB_ID/transcript?format=json" \
-H "Authorization: Bearer lm_live_YOUR_API_KEY"
The format parameter takes four values:
| format | what you get |
|---|---|
| json | timestamped segments with speaker labels |
| text | plain text |
| srt | subtitle file |
| vtt | WebVTT subtitle file |
JSON is paged, with a maximum of 500 segments per request. Speaker labels depend on your plan, because diarization is plan-gated.
The exact field names in the JSON are in the OpenAPI 3.1 spec that the developers page links to, so I do not copy them here.
Webhooks instead of polling
Webhooks need the Plus plan or higher. You can register up to 10 public HTTPS endpoints. The events are job.completed, job.failed, translation.completed and summary.completed.
curl -X POST https://libraryminds.com/api/v1/webhooks \
-H "Authorization: Bearer lm_live_YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://your-app.com/hooks/libraryminds",
"events": ["job.completed"]}'
The signing secret is returned once, when you create the webhook. Save it right away.
Every delivery carries an X-Libraryminds-Signature header. Its value is "sha256=" followed by the HMAC-SHA256 of the raw request body, using your secret. Check it before you trust anything in the body. Anyone can send a POST to a public URL, and the signature is how you know the request came from the API.
Node.js with Express (copied from the docs):
import crypto from "crypto";
app.post("/hooks/libraryminds", express.raw({ type: "application/json" }), (req, res) => {
const sig = req.header("X-Libraryminds-Signature") || "";
const expected = "sha256=" + crypto
.createHmac("sha256", process.env.LM_WEBHOOK_SECRET)
.update(req.body)
.digest("hex");
const ok = sig.length === expected.length &&
crypto.timingSafeEqual(Buffer.from(sig), Buffer.from(expected));
if (!ok) return res.status(401).end();
const event = JSON.parse(req.body.toString());
res.status(200).end();
});
Python (copied from the docs):
import hmac, hashlib
def verify(raw_body: bytes, signature: str, secret: str) -> bool:
expected = "sha256=" + hmac.new(
secret.encode(), raw_body, hashlib.sha256
).hexdigest()
return hmac.compare_digest(expected, signature)
The most common mistake is hashing the wrong thing. You must hash the raw bytes you received. If your framework parses the JSON first and you serialize it again, the bytes can change and every valid request will fail. That is why the Node example uses express.raw.
I ran both verify functions locally with a made-up secret and body. A correct signature passed, and a body with one changed character failed.
How delivery works:
The API tries up to five times, with backoff.
A stable delivery ID is sent on every retry, so you can ignore a repeat you already handled.
After three consecutive dead-lettered deliveries, the webhook is disabled until you enable it again.
So answer with 200 fast, then do the slow work after. The Node example above does exactly that.
Webhooks must reach a public HTTPS URL. For local work, put a tunnel in front of your dev server.
Errors you will hit
| status | code | usual cause |
|---|---|---|
| 401 | invalid_api_key, expired_api_key | wrong or old key |
| 402 | insufficient_credits | no transcription minutes left |
| 403 | insufficient_scope, plan_limit_reached | key scope or plan |
| 409 | idempotency_in_flight | same key still being processed |
| 422 | validation_failed, idempotency_key_reused | bad body, or key reused with a different body |
| 429 | rate_limited, concurrency_limit | too many requests or jobs at once |
| 5xx | internal_error, capacity_exceeded | server side, retry later |
Errors come in one shape, with a machine-readable code and a requestId:
{
"error": {
"code": "insufficient_credits",
"message": "Not enough transcription minutes remaining.",
"requestId": "req_9f2c1a4b8e7d6c5a4b3c2d1e"
}
}
Log the requestId. It is the first thing to send if you ask for help.
Scopes do not bypass plan gates. A key can have a scope and still get 403 if your plan does not include that feature.
Searching what you have transcribed
There is also POST /api/v1/search for semantic search over your library. It needs the search:read scope, and semantic search is on the Plus plan and above. I do not show the request body here, so check the docs for it.
The idea is simple. Each transcript is turned into vector embeddings, and results are ranked by meaning at paragraph level, not by how often a word appears. Searching "team morale" can find a paragraph about burnout or motivation, even if the words "team morale" never appear. Keyword search still exists for exact words and names. Use keywords for names and quotes, and semantic search for ideas.
API or the free tool?
If you only want one transcript now, the free YouTube Transcriber page works without a key. It reads an existing caption track, it is limited to 20 requests per 24 hours, and it has a Copy button but no file download. The API is for when code needs to do this for you, many times, and wait for the result.
Try it in 10 minutes
Create a free account and an API key with jobs:write, jobs:read and transcripts:read.
Run the Step 1 curl with a short video.
Check the job with Step 2 until it is done.
Fetch the transcript with format=text first, because it is the easiest to read.
When that works, register a webhook and add the signature check.
The full reference is the Libraryminds developer docs. You can create a free account, or use the no-key YouTube Transcriber.
If a step in the docs was unclear to you, tell me in the comments. I will fix the docs and this post.

