Tools you will use
Stack: From $55/moRiverside
Free trial · from $24Record remote interviews in local 4K quality, without the internet ruining the…
Read the reviewDescript
Free trial · from $16Edit video and audio by deleting words from the transcript, like in a document.
Read the reviewOpusClip
From $15Turns a long video into vertical clips ready for social, subtitles included.
Read the reviewTLDR: The mistake that ruins remote interviews is recording the video call: compressed audio and connection dropouts are baked in forever. Riverside records locally on each participant’s machine and returns separate tracks, Descript lets you edit the conversation by deleting text, and OpusClip pulls the social clips. All three have free tiers to test the pipeline on a real interview.
This guide is for anyone who interviews remotely and publishes the result: podcasts with guests, media outlets, companies recording with clients or experts. The remote interview is today’s best effort-to-value content format, because the guest brings the content and you only supply the process.
The rule that orders everything: quality is won in the recording, not in the edit. No plugin fixes video-call audio, which is why step two of this guide matters most.
1. Prepare the session and the guest
Eighty percent of remote interview problems are avoided with a five-line email beforehand. What to ask for is concrete: headphones (they eliminate echo, the one thing that genuinely cannot be fixed), a quiet room with the door closed, laptop plugged in and everything else closed in the browser.
Add two technical notes people do not know: the recording happens on their machine, so losing internet loses nothing, and when you finish they must not close the window until the upload completes. That second note prevents the most common disaster of this method.
And a two-minute test before really starting, with the guest talking while you watch levels. Catching a clipping mic there costs two minutes. Catching it while editing costs the interview.
2. Record locally with separate tracks in Riverside
Riverside solves the structural problem: during the session you see and hear each other over the normal connection, but each machine records locally at high quality. When you finish, those files upload. Quality depends on each person’s mic, not on the bandwidth of the worst-connected participant.
From that comes the advantage you hear in the result: separate video and audio tracks per participant. You can raise the quiet talker without touching the other, cut a cough without cutting the answer and cut between shots with judgment. That is exactly what separates a podcast that sounds like a studio from one that sounds like a call.
There is a free plan with two hours of multitrack recording, enough for a real test interview, and paid from $24 a month annually. The guest only has to open a browser link, which is the precondition for a busy guest saying yes.
3. Edit by text in Descript
An hour-long interview lands at forty publishable minutes, and finding what to cut on a timeline is painfully slow. Descript changes the gesture: it transcribes the conversation and you edit by deleting sentences from the text. The digression that goes nowhere, the question you repeated, the minute of noise before the guest warmed up.
The two automatic features that pay most here are filler word removal and silence compression. In an hour-long conversation that trims several minutes of padding and improves pacing more than any background music. Free plan with 60 minutes a month, paid from $16.
Use the transcript for two more things: the episode notes practically write themselves from it, and the subtitles for the next step’s clips come from there, so the fixes you make now propagate.
4. Pull clips with OpusClip and publish
A good interview yields weeks of content once sliced. OpusClip analyzes the long recording and proposes the hook fragments already vertical and subtitled, ranked by potential. Of ten proposals two or three are usually publishable, and that sifting takes minutes instead of an afternoon.
The selection criterion for an interview differs from a monologue: the clips that work contain a guest statement that stands alone, without needing your question. If the fragment starts with “yes, exactly”, it does not work however good what follows is.
With the full pipeline built, the process per interview lands at about two hours of your work: prep, recording, editing and clips. The publishing detail of each step is in the clips from long video guide, and if the guest speaks another language, the AI dubbing one. The whole sector lives in AI for content and media.
Common mistakes
Recording the video call. This is the original sin: compressed call audio cannot be fixed afterwards. Record locally or the result sounds like what it is.
Letting the guest close the window before files upload. It is the silliest way to lose a good interview. Warn them at the start and remind them at the end.
Not requiring headphones. Without them, each mic also records the other person’s voice through the speaker, and that echo survives even separate tracks.
Publishing the whole interview without slicing it. The full episode reaches people who already follow you. The clips bring new people. Publishing only the long version wastes 90% of the work.
Frequently asked questions
What happens if the guest’s internet drops?
Local recording continues: the live conversation is interrupted but their file keeps recording. When the connection returns, it uploads complete. That is the main reason to use this method.
How much does the pipeline cost?
You can test the whole thing free on all three tools’ free tiers. Paying, they total from about $55 a month, and can be adopted in phases: recording first, where the biggest gain sits, and the rest when volume demands it.
Does the guest need to install anything?
No, they join via a browser link. You only need to ask for headphones and that they leave the window open until the upload finishes.
Is it worth it if I only do one interview a month?
At one a month, the free tiers cover most of the pipeline. Subscriptions start paying off when you record weekly, or when audio quality is part of your pitch to guests.
The steps, in short
Prepare the session and the guest
Send short instructions on headphones, a quiet room and the browser, and test the setup before starting.
Record locally with separate tracks in Riverside
Let each machine record at high quality and wait for the files to upload before anyone closes the window.
Edit by text in Descript
Cut the conversation by deleting sentences from the transcript and strip filler words and silences.
Pull clips with OpusClip and publish
Turn the best moments into vertical pieces to feed social for weeks.
Related guides
How to run a flipped classroom with AI-made video lessons
Guide to the flipped classroom with AI: video lessons with Descript and ElevenLabs, comprehension…
Updated July 25, 2026How to run your clinic's social media with AI without stepping on red lines
Guide to healthcare social media with AI: content that builds trust without giving advice, the…
Updated July 25, 2026How to turn long video into social clips with AI
Guide to slicing long video into vertical clips with AI: moment selection with OpusClip, fine cuts…