> Canonical guide: https://developers.foxlight.ai/build/tutorials/video-studio/child/
> Contract snapshots: Skulk 2.0.0 (b0af39c79b6b7102b2478062904f1d7cc8619975); SDK 0.4.0 (31bb090b8f64689f87514e48385e22b6ad94a6c2). Check the installed runtime when versions differ.

# 3. Start and supervise the child

Your application begins as a supervised process. The host supplies startup identity, its manifest, configuration and a local socket. Your entrypoint must authenticate, answer the declared calls and release its resources when the host shuts it down.

Read `src/foxlight_video_studio/child.py` alongside this chapter. The complete [managed echo example](https://developers.foxlight.ai/build/examples/echo/) is the smaller version of the same contract.

## Read startup from the inherited descriptor

Studio's `main()` accepts `--config-fd`, uses `read_startup()` to read the bounded record, then starts its async loop. The installed process working directory owns its application data directory.

The host creates this startup record. You do not invent a token, node id or manifest to make a production child connect. You also do not put the token into a command-line argument, log or ordinary configuration field.

`Startup.node_id` identifies the capability installation. `transport_id` identifies the associated fabric node. `host_id` identifies the host endpoint. Similar-looking ids are not interchangeable.

## Authenticate and stay alive

This is Studio's actual managed loop:

```python title="studio:child.py:run"
async def run(startup: Startup, root: Path) -> None:
    """Serve the owner's IPC until shutdown; the worker runs beside it."""
    studio = Studio(startup, root)
    reader, writer = await asyncio.open_unix_connection(startup.socket)
    try:
        await write_message(
            writer,
            Hello(
                host_id=startup.host_id,
                transport_node_id=startup.transport_id,
                capability_node_id=startup.node_id,
                bundle_id=startup.manifest.bundle_id,
                bundle_version=startup.manifest.bundle_version,
                token=startup.token,
                manifest_sha256=manifest_digest(startup.manifest),
            ),
        )
        await studio.start()
        while True:
            message = await read_message(reader)
            if isinstance(message, Shutdown):
                return
            if isinstance(message, Health):
                await write_message(
                    writer,
                    Health(nonce=message.nonce, ready=True, surfaces=studio.surfaces()),
                )
            elif isinstance(message, Invoke):
                await write_message(writer, await studio.invoke(message))
    finally:
        await studio.stop()
        writer.close()
        await writer.wait_closed()
```

Walk it in order:

1. Construct application-owned state and services from the startup record.
2. Connect to the host-provided Unix socket.
3. Send `Hello` with the startup identities, token and digest of the installed manifest.
4. Start the application workers and HTTP surface.
5. Answer each `Health` challenge with the same nonce and current surface reports.
6. Dispatch each `Invoke`, then write one correlated `Result`.
7. On `Shutdown` or connection failure, stop owned work and close the socket in `finally`.

The SDK's `wire` helpers own bounded framing. You own the lifetime of the services you start. A working Python process alone does not mean a valid handshake occurred.

## Wire the durable services

In `Studio.__init__`, the render path is assembled in this order:

| Object               | Constructed from                                    | Responsibility                                               |
| -------------------- | --------------------------------------------------- | ------------------------------------------------------------ |
| `OperationJournal`   | `operations.sqlite3` + operation bounds             | Durable intent, status, backend reference and replay fences. |
| `OperationLogs`      | Protected log directory + bounds                    | Bounded, paged logs.                                         |
| `SkulkRenderAdapter` | HTTP client, asset/render directories and settings  | Submit, observe, collect and cancel a Skulk job.             |
| `OperationRunner`    | Journal, logs, adapter, scope and result schema     | One background worker for the render queue.                  |
| `OperationService`   | Same scope and stores, runner and planning callback | Dispatch the six operation verbs.                            |

Refinement gets its own journal, adapter, runner and service. Sharing the application process does not mean unrelated workflows should share an adapter or accidentally inspect each other's rows.

The journal is how a completed control call becomes durable work. The runner is how that work continues after the control call returns.

## Read the actual SDK construction

The constructor below creates both workflows. Start with the render section:
`OperationJournal` → `SkulkRenderAdapter` → `OperationRunner` →
`OperationService`. Follow `self.scope` through each constructor. The result
schema belongs to the runner because completion must be validated before it is
persisted, not only when a caller later asks for status.

Then compare refinement: it has a separate journal, adapter, runner and runtime
bound. Both workflows share the bounded log store while their scopes keep
records distinct. The HTTP server receives application handlers and file
stores; it does not own a second render runner.

<details>
<summary>Full Studio constructor — application and SDK resources</summary>

```python title="studio:child.py:Studio.__init__"
def __init__(
    self, startup: Startup, root: Path, api: SkulkApi | None = None
) -> None:
    self.startup = startup
    self.manifest = startup.manifest
    self.settings = (
        StudioSettings.model_validate(startup.configuration)
        if startup.configuration is not None
        else StudioSettings()
    )
    self.root = root
    secure_directory(root)
    self.api = api or SkulkApi(self.settings.skulk_api_url)
    declared = self.manifest.operations
    assert declared is not None
    # The operator's render timeout is the wall-clock bound the runner
    # cancels at; the manifest's bound is the ceiling it can never exceed.
    bounds = declared.model_copy(
        update={
            "max_runtime_seconds": min(
                self.settings.render_timeout_seconds, declared.max_runtime_seconds
            )
        }
    )
    self.scope = scope_for(startup.node_id, self.manifest, RENDER_ID)
    self.refine_scope = scope_for(startup.node_id, self.manifest, REFINE_ID)
    # Before either journal opens: bring records filed under an earlier
    # contract revision into the stable scope, so no take drops out of
    # the page when a release changes the contract.
    adopt_history(root / "operations.sqlite3", root / "logs", self.scope)
    adopt_history(root / "refinements.sqlite3", root / "logs", self.refine_scope)
    self.journal = OperationJournal(root / "operations.sqlite3", bounds)
    # The log directory is validated like the media directories: a link
    # here would let predictable log names read or write elsewhere.
    secure_directory(root / "logs")
    self.logs = OperationLogs(root / "logs", bounds)
    # The built web app, when one is packed; the server serves it first.
    self.webapp = WebApp()
    self.marks = TakeMarks(root / "takes.json")
    # Every finished render teaches its host's speed on its card; the
    # estimates are learned here rather than carried from other machines.
    self.timings = TimingLog(root / "timings.json")
    self.adapter = SkulkRenderAdapter(
        self.api,
        assets=root / "assets",
        renders=root / "renders",
        default_model=self.settings.default_model,
        keep_renders=self.settings.keep_renders,
        kept=lambda: self.marks.kept_jobs,
        timings=self.timings,
    )
    self.runner = OperationRunner(
        self.journal,
        self.logs,
        self.adapter,
        scope=self.scope,
        poll_seconds=5.0,
        adapter_seconds=float(
            min(ADAPTER_CALL_SECONDS, self.settings.render_timeout_seconds)
        ),
        result_schema=RenderResult.model_json_schema(),
    )
    self.service = OperationService(
        scope=self.scope,
        journal=self.journal,
        logs=self.logs,
        runner=self.runner,
        plan=self._plan_operation,
    )
    self.assets = AssetLibrary(root / "assets")
    # Refinements have their own journal and bound: the runner cancels at
    # the journal's runtime bound, and the refine bound is the operator's
    # own setting, not the render's.
    refine_bounds = declared.model_copy(
        update={
            "max_runtime_seconds": min(
                self.settings.refine_timeout_seconds, declared.max_runtime_seconds
            )
        }
    )
    self.refine_journal = OperationJournal(
        root / "refinements.sqlite3", refine_bounds
    )
    self.refine_adapter = ChatRefineAdapter(
        self.api,
        self.assets,
        root / "refinements",
        refine_model=self.settings.refine_model,
        timeout_seconds=float(self.settings.refine_timeout_seconds),
    )
    self.refine_runner = OperationRunner(
        self.refine_journal,
        self.logs,
        self.refine_adapter,
        scope=self.refine_scope,
        poll_seconds=REFINE_POLL_SECONDS,
        adapter_seconds=REFINE_ADAPTER_SECONDS,
        result_schema=RefineResult.model_json_schema(),
    )
    self.refine_service = OperationService(
        scope=self.refine_scope,
        journal=self.refine_journal,
        logs=self.logs,
        runner=self.refine_runner,
        plan=self._refine_plan,
    )
    self.backfill = ThumbnailBackfill(
        self.journal, self.adapter, self.marks, scope=self.scope
    )
    self._stills: asyncio.Task[int] | None = None
    # Set once the first still pass is done; a take listing waits for it.
    self._stills_settled = asyncio.Event()
    self._reach: asyncio.Task[None] | None = None
    self.reach_retry_seconds = REACH_RETRY_SECONDS
    self.reach_refresh_seconds = REACH_REFRESH_SECONDS
    self.server = StudioServer(
        local=self.local,
        assets=self.assets,
        renders=root / "renders",
        refinements=root / "refinements",
        webapp=self.webapp,
    )
```

</details>

## Give the journal a stable scope

```python title="studio:child.py:scope_for"
def scope_for(node_id: str, manifest: Manifest, descriptor_id: str) -> str:
    """The journal scope of one descriptor: installed node plus qualified identity.

    Not the descriptor's revision: that is a digest of its schemas, which a
    release changes whenever the contract grows a field, and a scope that
    moved with it hid every earlier take (``history.adopt_history``).
    """
    descriptor = next(d for d in manifest.descriptors if d.id == descriptor_id)
    return f"{node_id}/{descriptor.qualified_id}"
```

Studio uses the installed node plus qualified capability identity. A descriptor **revision** changes when a schema changes; making it the permanent history namespace would hide previous takes after an upgrade. Studio migrates earlier revision-based scopes before opening its journals.

For your application, keep installed identity stable across restarts. Plan data migration explicitly when you change durable keys. A new release should not make old work disappear merely because its descriptor grew a field.

## Dispatch is bounded

Studio maps a capability id to its corresponding handler. Render and refinement parse `OperationCall`; planning parses `RenderRequest`; assets and takes parse their own verb models.

`Studio.invoke()` budgets dispatch below the caller's remaining time and the manifest's 20-second call limit. It translates timeout and validation failures into typed SDK results and checks encoded reply size before writing it. Huge listings or raw media cannot be returned just because they fit in a Python object.

The main loop handles control calls sequentially. Long inference lives in `OperationRunner`, not in that loop. Blocking file work is moved off the async event loop. Otherwise an innocuous asset scan or render can prevent the process from answering health.

## Child health, surface health and render readiness

Studio answers child health while a model is missing. It separately reports whether the page can be reached and whether the fleet can render. These states serve different purposes:

- Child health: is the protocol endpoint responsive?
- Surface report: is this declared interface available at a usable URL?
- `video.readiness`: does the cluster currently have the model and engine capacity this application needs?

Restarting a healthy application repeatedly will not mount a missing model.

## Shut down the resources you own

Studio stops address-discovery and thumbnail tasks, stops its HTTP server, stops both operation runners and closes its HTTP client. Its durable rows and saved files remain. Stopping the child worker is not a license to kill an inference process the application does not own.

Next: [resolve plans and readiness](https://developers.foxlight.ai/build/tutorials/video-studio/planning/).
