What Ships to the Pod#
Kinetic uploads two artifacts to Cloud Storage for every job. The pod downloads both artifacts and rebuilds your project in a workspace directory.
payload.pkl— your function, the objects that it closes over, its arguments, the environment variables that you capture, and a fingerprint of your client toolchain. Kinetic serializes all of these withcloudpickle.context.zip— a snapshot of your project source, rooted at the package root. Kinetic writes the packaging plan into the same archive, at the reserved path.kinetic/plan.json.
Kinetic reads one dependency file too: a requirements.txt or a
pyproject.toml. In bundled mode Kinetic puts that content into the
image build. In prebuilt mode Kinetic uploads the content beside the two
artifacts, and the pod installs it at startup. See
Dependencies.
This page is the exact contract for the artifacts and their destinations.
Read it first when you debug a ModuleNotFoundError, a
FileNotFoundError, or a job that uploads gigabytes. Then read
Troubleshooting.
The package root#
The package root is the directory that Kinetic archives. Kinetic resolves it at submit time, in three steps.
Step 1 — find the entry directory. Kinetic finds the module that
defines the decorated function, and takes the directory of that module
file. A notebook cell, a REPL, and python -c define no module file. In
those cases Kinetic uses your current working directory instead.
Step 2 — escape the package. Kinetic walks up from the entry
directory for as long as each directory holds an __init__.py file.
This step makes from trainer.model import ... resolvable on the pod,
because the root ends above trainer/, and not inside it. This walk also
stops at your home directory and at the root of the file system.
Step 3 — walk up to a project marker. Kinetic then walks up to the nearest directory that holds one of these markers:
pyproject.tomlrequirements.txtsetup.pysetup.cfg.git(a directory, or a file — a git worktree uses a file)
The walk stops at your home directory and at the root of the file system. Kinetic never adopts either one as the package root, unless step 2 already ended there. Kinetic keeps the directory from step 2 when it finds no marker.
Override. Set KINETIC_PACKAGE_ROOT to pin the root:
export KINETIC_PACKAGE_ROOT=/home/me/monorepo/services/trainer
The override replaces all detection. Kinetic expands a leading ~ in the
value. Kinetic then validates the value, and raises a ValueError at
submit time in two cases:
The value does not name an existing directory.
The value is neither the entry directory nor a parent of it.
Tip
Two habits keep the root predictable:
Keep a
pyproject.toml, arequirements.txt, or a git repository at the top of the tree that you want to ship.Keep large data out of that tree. As an alternative, wrap the data in
kinetic.Data(...), which Kinetic excludes fromcontext.zipautomatically.
Worked examples#
Layout |
Decorated function in |
Package root |
Why |
|---|---|---|---|
|
|
|
The entry directory already holds a marker. |
|
|
|
Step 2 escapes the package. Step 3 finds the marker. |
|
|
|
|
|
|
|
No marker exists. Kinetic keeps the entry directory, and never |
A Jupyter notebook in |
the notebook cell |
|
The cell has no |
What Kinetic excludes#
Kinetic always excludes these two directory names, at any depth:
.git, __pycache__
Kinetic also excludes these names at any depth, unless you turn the default exclusions off:
.venv, venv, node_modules, .tox, .mypy_cache, .ruff_cache,
.pytest_cache, .ipynb_checkpoints, .DS_Store
Kinetic excludes a local path that you wrap in kinetic.Data(...) as
well, when that path lies inside the package root. Those bytes travel
through the content-addressed cache for data instead, so Kinetic does not
upload them two times. See Data.
To turn the default exclusions off, set this variable:
export KINETIC_NO_DEFAULT_EXCLUDES=1
The value must be exactly 1. .git and __pycache__ stay excluded.
.kineticignore#
Put a .kineticignore file at the package root to exclude more paths.
Kinetic reads this file at the package root only. Write one pattern per
line, in fnmatch syntax. A line that starts with # is a comment.
Kinetic reads a # in any other position as part of the pattern. A
trailing / restricts the pattern to directories.
# .kineticignore
checkpoints/
*.ckpt
scratch_*.py
docs/_build/
Kinetic matches every pattern two times:
Against the path relative to the package root, such as
docs/_build.Against the last name in that path, such as
_build.
The second match applies at any depth. *.ckpt therefore excludes a
checkpoint file anywhere in the tree, and checkpoints/ excludes every
directory with that name.
Secret files#
Kinetic does not exclude secret files by default, because some projects need them on the pod. Kinetic logs a warning instead when a file name in the archive matches one of these patterns:
.env**.pemid_rsa*
The warning names every matched file. If you did not intend to send one of the files, do these steps:
Add the file to
.kineticignore.Submit the job again.
Size warnings#
Kinetic logs the archive size on every submit. Kinetic logs a warning
above 100 MB, and lists the five largest files in the archive. Change the
threshold with KINETIC_CONTEXT_SIZE_WARN_MB.
The pickled payload has a separate threshold of 50 MB
(KINETIC_PAYLOAD_SIZE_WARN_MB). A large payload almost always means one
module-level global that Kinetic captured by value. cloudpickle
serializes the objects that your function references. A module-level
DF = pd.read_parquet(...) that your function reads therefore lands
inside payload.pkl on every submit. Pass large data as
kinetic.Data(...), or load it inside the function.
Set either threshold to 0 to turn that warning off. Kinetic ignores a
value that is not a number, and logs a warning about the bad value.
Fidelity of the archive#
context.zip is a faithful snapshot, with these explicit rules:
Kinetic follows a symlinked directory, and guards against a cycle. Kinetic archives each real directory one time only, and names every skipped directory in a warning.
Kinetic skips a broken symlink and an unreadable file, and logs a warning for each one. One bad file never stops a submit.
Kinetic stores the POSIX mode of each file. The runner restores the read, write, and execute bits on the pod, for an archive that a Unix client built. The runner never restores a setuid, setgid, or sticky bit.
Kinetic archives an empty directory, and the runner restores it.
Kinetic clamps a timestamp that falls outside the range of the ZIP format, which is the year 1980 to the year 2107.
What happens on the pod#
The pod downloads both artifacts and verifies their SHA-256 hashes. The runner then does the following:
The runner extracts
context.zipinto the workspace directory.The runner rebuilds
sys.pathfrom the packaging plan. The workspace root goes first. Each clientsys.pathentry that lived under the package root follows it, at the matching position inside the workspace. This step makes asrc/layout work without aPYTHONPATHon the pod.The runner changes the working directory to the workspace directory that matches your client working directory. The runner uses the workspace root when your client working directory was outside the package root.
The runner unpickles
payload.pkland calls your function.
The runner does not replicate a client sys.path entry that points at
site-packages, or at any directory outside the package root. The image
provides those packages.
Step 3 has one practical consequence: a relative path behaves the same
on the pod as it does on your machine. open("configs/train.yaml")
works remotely if it worked locally from the same directory. The file
must be inside the package root, and it must not be excluded.
An absolute client path is not a supported way to read a shipped file. The runner does create one symbolic link at the path of your client working directory. The link points at the workspace root. The link exists so the debugger can map source files. Use relative paths for portable access to files.
The pod environment also holds KINETIC_OUTPUT_DIR. The value is a Cloud
Storage URI, such as gs://{bucket}/outputs/{job_id}. The pod destroys
the workspace at the end of the job. Write everything that you want to
keep to that Cloud Storage location. See
Checkpointing.
How imports resolve remotely#
Kinetic uses two mechanisms. The mechanism that applies explains almost every remote import failure.
Kinetic ships your first-party code by value. At submit time Kinetic
inspects sys.modules. Kinetic registers each module whose file lives
under the package root with cloudpickle.register_pickle_by_value.
cloudpickle then serializes the functions and classes of those modules
into payload.pkl in full. The pod therefore does not import your
modules to unpickle your job. A helper in trainer/utils.py travels
inside the payload. A module-level import kinetic in one of your modules
does not have to succeed at unpickle time.
Kinetic never registers these two modules by value:
The
kineticpackage itself, even when you work inside a checkout of Kinetic.The
__main__module.cloudpicklealready ships a__main__function by value.
The pod resolves third-party imports. cloudpickle pickles anything
outside the package root by reference: numpy, keras, and your private
packages. The image must provide those packages, and that is the purpose
of the dependency file.
The workspace resolves runtime imports. An import statement inside
your function runs at call time, against the pod sys.path. That path
starts with the workspace. The pod finds a first-party module that you
import lazily in the extracted context. The image must provide a
third-party package that you import lazily.
Note
Identity limit of shipping by value. A class that travels by value
is a different type object from the same class that the pod imports. An
isinstance check across the two paths can return False. A class
defined in __main__ always behaves this way. Compare by attribute, or by
name, when you must cross that boundary.
Arguments: which types Kinetic preserves#
Kinetic pickles the arguments and the keyword arguments with the
function. Kinetic walks them for one purpose only: to replace each
Data(...) object with a reference. Kinetic scans the arguments on every
submit. Kinetic rebuilds the containers only when the call holds a Data
object. The pod also skips its own walk when the payload holds no
reference.
The walk keeps types and object identity:
Kinetic rebuilds a
list, and a list subclass, as the same type.Kinetic rebuilds a
tupleand aNamedTupleas the same type, with the fields intact. Bothtyping.NamedTupleandcollections.namedtuplework.Kinetic passes a
setand afrozensetthrough unchanged, because neither one can hold aDataobject.Kinetic preserves a
dictsubclass.OrderedDictkeeps its order,defaultdictkeeps itsdefault_factory, andCounterkeeps its type.Kinetic preserves aliasing. If one object appears two times in your arguments, your function receives the same object two times (
out[0] is out[1]).Kinetic pickles everything else whole, and it arrives unchanged.
If Kinetic cannot rebuild a container subclass, Kinetic uses the plain built-in type instead, and logs a warning. Kinetic does not fail the job.
Kinetic rejects three argument shapes at submit time. Each message names the position of the argument at fault:
A
Dataobject inside asetor afrozenset. The replacement reference is a dict, and a dict is not hashable.A
Dataobject used as adictkey.A self-referential structure whose cycle runs through a tuple, a set, or a frozenset. A cycle through a list or a dict arrives on the pod unchanged. Kinetic reports the cycle only when the call also holds a
Dataobject, because only then does Kinetic rebuild the containers.
Kinetic does not find a Data object inside the attributes of a custom
object, because Kinetic walks plain containers only. Kinetic uploads
nothing for that object, and the pod hands your function the original
Data instance, which points at a path on your machine. Pass a Data
object as its own argument, or inside a list, a tuple, or a dict.
When pickling fails, Kinetic bisects the payload and names the component
at fault. One example message is
kinetic could not serialize argument 2 (type socket): .... Kinetic
counts positional arguments from 0, and identifies a keyword argument
by name. You therefore never get a bare PicklingError from inside
cloudpickle.
Environment variables#
capture_env_vars sends local environment variables to the pod. Kinetic
always accepts an exact name. A name that ends with * is a prefix
pattern: "WANDB_*" matches every name that starts with WANDB_, and
"*" matches every name. A * in any other position is not a wildcard.
A prefix pattern never matches the names in this blocklist. The pod applies your values over its own environment, and these names break the pod runtime:
PATH, HOME, PYTHONPATH, LD_LIBRARY_PATH, LD_PRELOAD,
VIRTUAL_ENV, CONDA_PREFIX, CONDA_DEFAULT_ENV, SHELL, TMPDIR,
TEMP, TMP, HOSTNAME, USER, LOGNAME, SSH_AUTH_SOCK,
KUBERNETES_SERVICE_HOST, KERAS_BACKEND
You can still forward any of these names. List the name exactly:
capture_env_vars=["KERAS_BACKEND"] works, and a prefix pattern that
covers it does not.
Kinetic logs the names that it captured, and never the values. Kinetic
logs one more warning when a captured name contains TOKEN, SECRET,
KEY, PASSWORD, or CREDENTIAL, in any letter case. Kinetic stores
those values inside payload.pkl in the job bucket, and every job pod in
the cluster can read that bucket. This warning is informational, and it
appears even for a name that you listed exactly. If you intend to send the
credential, you need no action. See
Forward Environment Variables and
Security.
Notebooks and REPLs#
A function defined in a Jupyter cell, an IPython session, or python -c
has no source file. Kinetic detects this condition. Kinetic then uses your
current working directory as the entry directory, and logs the directory
that it chose. Steps 2 and 3 then run as usual, so a project marker above
your current directory can move the root further up the tree. Change into
your project directory before you submit.
A notebook function has no importable module, so cloudpickle always
pickles it by value. That behavior is the one that you want. A helper
from another cell travels with the function. A helper in a .py file
beside the notebook travels in context.zip.
Matching your local environment to the pod#
Pickled code objects are not portable across Python minor versions. A function pickled on 3.12 does not unpickle on 3.11.
Bundled mode (the default) builds an image from
python:{your minor version}-slim. The build installskeras,cloudpickle,google-cloud-storage, JAX for your accelerator category, andkeras-kineticpinned to your client version. The pod Python therefore always matches your client. Use this mode if you do not want to think about the question.Prebuilt mode pulls a base image that you publish with
kinetic build-image. Kinetic requests the image tag that matches your client Kinetic version. Build that image with the Python minor version of the client that submits the job.Custom image mode makes you responsible for the Python version. See the image requirements in Execution Modes.
Every payload carries a fingerprint of the client: the Python version,
the cloudpickle version, and the Kinetic version. The runner compares
the Python version and the cloudpickle version against the pod, and logs
a warning about a difference. The error that result() raises names both
sides, such as client Python 3.12.2 / pod Python 3.11.9. The pod log
holds a skew warning only when the payload unpickled correctly.
Environment knobs#
Variable |
Default |
Effect |
|---|---|---|
|
(unset) |
Pin the package root. The value must name an existing directory, and must be the entry directory or a parent of it. |
|
(unset) |
Set it to |
|
|
The |
|
|
The |