--- name: burla-usage version: 1.4.0 description: Execute Python code in the user's cloud using Burla, guides correct remote_parallel_map use. Use whenever a Python task would be slow on one machine (many files, many API calls, big data processing, parameter sweeps, simulations, model inference, anything embarrassingly parallel) and the user is signed in to AWS, Google Cloud, or Azure, even if Burla is not installed yet, or for work on anything that already calls remote_parallel_map. Covers checking for and setting up Burla, every argument, nested calls, pipelines, CLI control and debugging, and dashboard URLs. --- # Using Burla ## What Burla does Burla runs ordinary Python on many cloud machines at once. You give `remote_parallel_map` one Python function and a list of inputs; Burla boots or reuses VMs in the user's cloud account, installs the same packages the user has locally, runs the function on every input at the same time (one call per CPU, spread across every machine), streams prints and errors back, and returns the list of results. Nothing else to deploy: no containers to build, no cluster to configure, no job framework to learn. ```python from burla import remote_parallel_map def process(path): return analyze(path) results = remote_parallel_map(process, paths) ``` The whole run is one job, visible live on the dashboard. Reach for it whenever Python needs more machines, more time, or more cores than the laptop has. ## Setup When a task looks like a fit for Burla, check whether it is ready before writing Burla code, and set it up if it is not. The user does not need a Burla account, `burla login`, or `burla deploy`: Burla runs its dashboard on this machine and boots VMs with the user's own cloud CLI credentials. Being signed in to one cloud with permission to create VMs is the only requirement. 1. **Find a signed-in cloud.** Run the check for each CLI that is installed (`which aws gcloud az`); a zero exit code with output means signed in: ```bash aws sts get-caller-identity gcloud config get-value project && gcloud auth application-default print-access-token az account show ``` GCP needs both a project (the first command exits 0 even when it prints nothing or `(unset)`, which means no project: `gcloud config set project `) and application default credentials (`gcloud auth application-default login`). If no cloud is signed in, stop and ask the user to sign in (`aws sso login` or `aws configure`, `gcloud auth login` plus `gcloud auth application-default login`, or `az login`). Do not sign in for them or fall back to running the task locally without saying so. 2. **Install Burla** into the same Python environment that will run the task (Python 3.11+). Burla copies that environment's packages onto the VMs, so installing into a different interpreter breaks imports remotely: ```bash python -c "import burla" 2>/dev/null || python -m pip install burla ``` 3. **Pick the cloud** only when more than one CLI is installed. Burla picks the only installed CLI automatically, but with several it asks interactively, which an agent's shell cannot answer. Set it to the cloud that passed step 1 (ask the user if several did): ```bash burla config set cloud aws # or gcp, azure ``` 4. **Start the dashboard** before the first job and get its URL: ```bash python -c "import burla; print(burla.get_cluster_dashboard_url())" ``` This starts Burla's dashboard as a background process on this machine (no browser opens) and prints its URL, e.g. `http://127.0.0.1:5001`. It keeps running after jobs finish, so job pages stay viewable. Without this step `remote_parallel_map` still works, but it starts its own dashboard and shuts it down when the job ends, so the job URL goes dead. 5. **Run the task** with `remote_parallel_map`, then give the user the job URL, saying the dashboard is running locally on their machine. Use `grow=True` for any job big enough to justify Burla, since a fresh cluster has no idle VMs. VMs cost the user money and delete themselves after going idle. If a run fails at startup, read the error: Burla's setup errors name the exact command that fixes them (e.g. an expired AWS SSO session or an unset gcloud project). Run it or ask the user to, then retry. ## Updates Run this once, before the first use of this skill in each agent session. Set `SKILL_DIR` to the absolute directory containing this `SKILL.md`. If `SKILL_DIR` is not itself a Git repository, skip the update check. This is the monorepo source copy or a manually copied installation. For a Git installation, verify that `origin` is `https://github.com/Burla-Cloud/burla-skill.git` or `git@github.com:Burla-Cloud/burla-skill.git`. If it is anything else, skip the update and tell the user. Then fetch the public repository without changing local files: ```bash git -C "$SKILL_DIR" fetch --quiet origin main ``` If `HEAD` equals `origin/main`, continue. If `HEAD` is an ancestor of `origin/main`, read the version from `origin/main:SKILL.md`. Continue with the installed version when the remote version is unchanged; only offer an update when the version increased. If `~/.config/burla-skill/auto-update` contains `true`, run: ```bash git -C "$SKILL_DIR" pull --ff-only origin main ``` Otherwise ask the user: "Burla usage skill v{new} is available (installed: v{old}). Update now?" Offer **Update now**, **Always keep updated**, and **Not now**. Only pull after **Update now**. For **Always keep updated**, write `true` to `~/.config/burla-skill/auto-update`, then pull. For **Not now**, continue with the installed version. Before pulling, require a clean working tree and verify that `HEAD` is an ancestor of `origin/main`. Never reset, overwrite local changes, or pull a divergent branch. Report that the installation was modified and continue with the installed version. After a successful update, re-read `SKILL.md` and follow the new version. ## Vocabulary One `remote_parallel_map` call is one job; each input is one function call, identified by its 0-based input index. Exceptions re-raise on the client with `exc.burla_input_index` set to the input that failed. ## Three rules that always apply 1. **Always give the user the job URL.** Every job has a live dashboard page at `{dashboard_url}/jobs/{job_id}`. Whenever you start a job for the user, find its id (`burla jobs list --limit 1`, most recent first) and print the full clickable URL so they can watch progress, logs, and metrics. Get the dashboard URL from `burla auth status` (`data.head_url`) or `python -c "import burla; print(burla.get_cluster_dashboard_url())"`. Never ask the user to run `burla dashboard` (it is for humans and opens a browser); the `get_cluster_dashboard_url()` call above starts the dashboard on this machine if needed and just prints its URL. In a burla dev worktree use that worktree's own cluster (`make cluster-info`, see the burla-parallel-dev skill). 2. **Use the `burla` CLI for cluster control and inspection.** Never poke Burla's HTTP endpoints or guess at state: the CLI covers cluster lifecycle, job inspection, per-call debug logs, node logs, settings, and usage. See [the CLI section](#the-burla-cli) below. 3. **A pipeline is one parent job.** Any script or workflow containing more than one `remote_parallel_map` call should run inside a single parent `remote_parallel_map` call. See [Pipelines](#pipelines-and-nested-calls). ## Every argument, and when to use it `function_` (required): any Python callable taking one input item. Pickled with everything it references, which must total under 100MB; pass big objects as inputs or load them inside the function instead of referencing them from an enclosing scope. Burla replicates the client's whole Python environment on workers (every installed package at its exact version, plus local module source), so keep imports at module scope and never move an import inside the function to "make it picklable". `inputs` (required): a list of pickleable objects, one per function call. Tuples are unpacked into positional arguments: `inputs=[(1, 2)]` calls `function_(1, 2)`; to pass a tuple as one argument, nest it: `[((1, 2),)]`. Submit one input per logical unit of work and let Burla schedule: never pre-batch inputs to match a core count, extra inputs queue and rebalance across nodes automatically. `func_cpu` (default `"dynamic"`): CPUs per call. Dynamic starts at one call per CPU and gives calls more CPU by lowering parallelism on nodes where calls measurably wait for a core. Pass an integer only when a call needs a guaranteed core count (e.g. a multithreaded library). `func_ram` (default `"dynamic"`): RAM in GB per call. Dynamic lowers parallelism on nodes where workers run out of memory. Pass an integer when you know the per-call footprint and want it reserved up front instead of discovered through OOM retries. `func_gpu` (default `None`): one GPU per call. One of `"T4"` (AWS only), `"A100"` / `"A100_40G"`, `"A100_80G"`, `"H100"` / `"H100_80G"`. Only nodes on a matching GPU family are eligible; combine with `grow=True` to boot GPU nodes when none are idle. `image` (default `None`): restrict the job to nodes running this container image; with `grow=True`, new nodes run it. Do not build an image just for pip packages, environment replication already installs those. Use an image for system-level dependencies: apt packages, CUDA, preinstalled tools or data. With `grow=True` and no image, new nodes run stock `python:3.X` matching the client's Python version. `grow` (default `False`): boot additional nodes (up to 2560 CPUs) so the job finishes as fast as possible. Use for any large job; leave off for small jobs that should just use idle nodes. `max_parallelism` (default: number of inputs): hard cap on concurrent calls. Without it, nodes may run more calls than vCPUs when CPU, memory, disk, and network all show headroom. Use it to protect rate-limited APIs or shared resources like a database, not to manage memory (that is `func_ram`). `detach` (default `False`): the job keeps running on the cluster if the local process exits, once inputs finish uploading (wait for the "Done uploading inputs!" message before killing the process). Requires a deployed cluster (`burla deploy`); a dashboard running locally cannot outlive its own process and raises `DetachRequiresDeployedCluster`. Use for long jobs you start and walk away from, then monitor with `burla jobs watch JOB_ID`. `generator` (default `False`): return a generator yielding results as they are produced instead of a list at the end. Use to pipeline downstream work, to show progress, or when all results at once would not fit in client memory. `raise_errors` (default `True`): by default the first exception in any call fails the whole job and re-raises on the client. Set `False` for fault-tolerant batch runs: failed calls return their exception object in the results (with remote traceback and `burla_input_index` attached), every other call runs to completion, and failures stay visible in the dashboard and `burla jobs errors`. `spinner` (default `True`): terminal status indicator. Auto-disabled when stdout is not a TTY; set `False` explicitly when you want clean parseable output. `region` (default `None`): pin the job to one cloud region (e.g. `"us-east-2"`), for data locality or compliance. Only idle nodes already in that region are eligible, and nodes booted for the job boot there. Default: idle nodes anywhere, new nodes in the cluster-settings region. `disk_gb` (default `None`): boot disk size for nodes booted for this job. Idle nodes are eligible regardless of their disk size. Default: the cluster-settings disk size. **Returns**: a list of whatever `function_` returned, in no particular order. If order or attribution matters, include the key in the return value (e.g. return `(input_id, result)`); never assume results align with input order. ## Pipelines and nested calls `remote_parallel_map` works inside a function that is itself running on the cluster: workers carry credentials, the nested call resolves the same cluster, and its job is automatically tagged with the enclosing job as its parent. The dashboard nests child jobs under the parent (structure graph and breadcrumb on the job page), and a parent that fails or is canceled cancels its running children. So: any pipeline, i.e. anything with multiple `remote_parallel_map` calls, should run inside one parent call. Write a driver function containing the stage calls and submit it as a single-input job. Burla detects the nesting and displays the whole pipeline as one job in the UI instead of several unrelated ones. Prefer `detach=True` on the parent when a deployed cluster is available so the pipeline survives the client disconnecting, but it is not required. ```python from burla import remote_parallel_map def align(sample): ... def call_variants(alignment): ... def pipeline(samples): alignments = remote_parallel_map(align, samples, grow=True) return remote_parallel_map(call_variants, alignments) results = remote_parallel_map(pipeline, [(samples,)], detach=True)[0] ``` Nested calls are also how one call fans out: e.g. a per-file job that spawns a per-record job for its own file. ## The burla CLI Management commands are non-interactive and print JSON (NDJSON for streams), built for exactly this kind of agent use. Full reference: [Burla management CLI](https://github.com/Burla-Cloud/burla/blob/main/docs/agent-cli.md) or `burla --help`. Control the cluster: ```text burla cluster status | start | restart | stop | watch burla settings show burla settings update --quantity N --machine-type TYPE --region REGION ... ``` Inspect jobs and debug failures: ```text burla jobs list --status running # find job ids, most recent first burla jobs show JOB_ID # status, counts, timing burla jobs watch JOB_ID # stream until it finishes burla jobs cancel JOB_ID burla jobs errors JOB_ID # failures grouped by traceback burla jobs calls list JOB_ID --failed-only burla jobs calls logs JOB_ID INPUT_INDEX # everything one call printed burla jobs metrics JOB_ID # utilization time series burla nodes list / show / logs NODE_ID # node-level (VM) logs ``` When a job fails, do not guess: `burla jobs errors JOB_ID` for the grouped tracebacks, then `burla jobs calls logs JOB_ID INPUT_INDEX` for the failing call's output. Target a specific cluster with `--head URL` or `BURLA_CLUSTER_DASHBOARD_URL`; check what the CLI is talking to with `burla auth status`. ## Common mistakes - Assuming results are ordered. They are not; carry keys in return values. - Forgetting the URL. The user should always get `{dashboard_url}/jobs/{job_id}`. - Running a multi-stage pipeline as sibling top-level jobs instead of inside one parent call. - Building a container image just to pip-install packages. - Moving imports inside the function; module-scope imports work. - Referencing large arrays or dataframes from an enclosing scope instead of passing them as inputs (functions over 100MB pickled are rejected). - Pre-batching inputs to core count instead of one input per unit of work. - Using `max_parallelism` for memory problems instead of `func_ram`. - Expecting `detach=True` to work against a locally hosted dashboard.