6. Add prompt refinement
Studio also helps an operator rewrite a shot description for the selected video workflow. It calls a ready chat model through Skulk™. This is a second application workflow, not an extension of the video engine.
Prepare the rewrite
ChatRefineAdapter.prepare validates RefineRequest, derives the reference mode, chooses model-specific guidance and resolves a ready chat model. Asset labels help describe the inputs; the chat request does not acquire unrestricted access to the asset directory.
The plan digest binds the request, chosen chat model, guide and labels. Preparation checks the available context budget and chooses a bounded output allowance. A prompt rewrite must leave enough room for the complete structured answer.
Unlike render plans, Studio's refinement OperationPlan does not return an effective input. A caller keeps its original refinement request and adds the returned digest when starting. Do not assume every service returns a ready-to-send input object merely because rendering does.
Use the same SDK, with a different backend
async def submit(self, operation_id: str, payload: dict[str, JsonValue]) -> str:
"""Start the rewrite; the operation id is the reference."""
prepared = await self.prepare(payload)
self._jobs[operation_id] = _Job(
task=asyncio.create_task(self._run(operation_id, prepared)),
prepared=prepared,
)
return operation_id
The adapter starts an owned asyncio task and uses the operation id as its reference. The SDK still records the durable operation, enforces bounds and answers status calls. The backend work here is an in-process chat request rather than an independently addressable video job.
That difference matters on restart: the operation journal survives, but this asyncio task does not. observe reports a missing task as failed and tells the caller to start again. Durable bookkeeping does not make every underlying computation resumable.
Inspect the generated result
async def _run(self, operation_id: str, prepared: Prepared) -> RefineResult:
started = time.monotonic()
async with asyncio.timeout(self.timeout_seconds):
reply = await self.api.chat_completion(
prepared.model,
prepared.messages,
max_tokens=prepared.max_tokens,
temperature=TEMPERATURE,
timeout_seconds=self.timeout_seconds,
)
if reply.finish_reason not in (None, "stop"):
# A reply cut by the token budget or a filter is not the prompt
# the guide asked for, whatever sections it happens to contain.
raise ReplyUnusable("the chat model stopped before finishing the rewrite")
prompt, warnings = parse_reply(reply.content, prepared.guide)
# The whole prompt lives in its own file; the journal row carries as
# much of it as fits its bound, and says when that is less.
path = self.refinements / f"{operation_id}.txt"
await asyncio.to_thread(self._keep, path, prompt)
result = RefineResult(
prompt=prompt,
prompt_path=str(path),
guide=prepared.guide,
guide_sha256=guide_sha256(prepared.guide),
model=prepared.model,
mode=prepared.mode,
labels=tuple(prepared.labels),
input_tokens=reply.input_tokens,
output_tokens=reply.output_tokens,
elapsed_seconds=round(time.monotonic() - started, 3),
warnings=tuple(warnings[:8]),
)
return fit_result(result)
This function demonstrates a complete application-owned transformation:
- Apply a timeout to the model call.
- Supply explicit
max_tokensand temperature. - Refuse a reply stopped early by a token budget or filter.
- Parse the required guide structure.
- Store the full rewritten prompt in a protected file.
- Return provenance and a bounded result.
The result records the model, guide identity and digest, mode, labels, token counts, elapsed time and warnings. If the full prompt exceeds the journal result budget, Studio returns a clipped representation with warnings and keeps the complete file separately. The HTTP refinement route can retrieve that full text.
Cancellation cancels the owned chat task; subsequent observation reports cancellation. Exceptions become bounded, safe failure reasons. Raw upstream errors can contain deployment details, so do not copy them straight into operator-visible journal fields.
Keep the operator in control
The browser presents the rewritten prompt for review. Accepting it updates the shot description; it does not silently launch a render. Planning still happens against the current video card and placement.
For your own application, decide whether model-assisted rewriting is a recommendation, an automatic transform or a governed action. Make that choice explicit in the contract and interface. A model's confident output is not proof that its suggested parameters are supported by the next service.
Next: assets and saved takes.