- tests are separated into
tests/unitwhich are quick to run with mocks, andtests/integrationwhich are heavyweight multi-process tests- when adding new functionality, try to add both unit tests and integration tests
- there are additionally
tests/largeE2E-- you are not expected to run or modify these during regular work, those are for releases
- project is managed by
uv-- utilize that for running any python-related subcommands likeuv run pytestoruv run ty - utilize
justfor command running --just valis the "all typechecking and testing". Always run this prior to a commit, as well as utilizing pre-commit with prek- during development, utilize granular validation -- first, run type checking, then unit tests for the code you have created or changed, then all integration tests
- don't change formatting on a whim. When you notice a bug or breach of guidelines that is not related or affecting your current task, ignore it.
- when you are creating a new package in backend/packages, initialize it with uv, add it to the backend/pyproject.toml workspace listing, and create there a basic justfile with the
valrecipe. When filling thejustfileand thepyproject.toml, use the existing packages as templates
- always use type annotations -- it is enforced
- when working with a package with insufficient typing coverage like sqlalchemy, use
ty:ignorecomment - when
tyis not powerful enough, usety:ignore - use
typing.castwhen the code logic is implicitly erasing the type information - do not use annotation-as-string unless necessary -- that is, prefer
f: TypeExpressioninstead off: "TypeExpression"
- when working with a package with insufficient typing coverage like sqlalchemy, use
- prioritize using pydantic BaseModel or dataclasses.dataclass object for capturing contracts and interfaces
- if an object would cross language/host boundaries go with pydantic to have json serde, otherwise dataclass. Do not use NamedTuples
- when using pydantic, use
FiabBaseModelfromforecastbox.utility.pydantic(orFiabCoreBaseModelfromfiab_core.pydantic_utilsin fiab-core) instead ofpydantic.BaseModeldirectly, unless the model requires dynamic field handling (e.g.,extra="allow"for JSON Schema types). These base models setextra="forbid"to catch misconfigured constructors early. If you need the dynamic model handling, mark it clearly with a comment. - ideally keep them plain, stateless, frozen, without functions -- we end up serializing those objects often over to other python processes or different languages
- having a few methods that provide convenience views, ser/de, validation, conversion does not necessarilly hurt -- its primarily about keeping the codebase generally functional and data-oriented rather than object-oriented.
- for simple immutable data transfer objects, use
@dataclass(frozen=True, eq=True, slots=True)directly for best type checker support -- provides immutability, hashability, and memory efficiency via slots. We seteq=Trueexplicitly, despite being a default, for clarity. - a convenience decorator
frozendcexists inforecastbox.utility.structuralbut direct decorator syntax is preferred for type safety - when using a primitive type in a semantically restricted context, utilize typing.NewType -- for example, dont do
user_id: strbutUserId = typing.NewType("UserId", str); user_id: UserId, because not every string is a valid UserId. This prevents a mixture of ids and gives a stronger type validity
- use comments sparingly, for non-obvious code only. Add docstrings to functions called from other modules only. When adding docstring, use compact style -- dont separate out Args and Returns, describe everything in one or two paragraphs. Do not make two spaces after a dot.
- when making a code change, do not refer to the previous state in the docstring or the comment. In other words, do not put to docstring texts like "this function replaces the original function". This belongs to commit message, not to the code!
- all imports belong to top level of the file, dont import inside function definitions unless necessiated by runtime (and in that case, always add explanatory comment at the site)
- dont alias in imports unless there is a name collision, or unless its a standard shortcut:
datetime as dt,multiprocessing as mp,numpy as np,xarray as xr,earthkit.data as ekd - never use python keywords and builtins as variable names -- for example, don't use
idvariable, preferid_<something>orid_
When adding new code, make sure you place it in the right submodule:
- routes: all backend routes are declared here. There is autodiscovery mechanism, do not make submodules here; only declare routes in
routes/*.pyfiles.- when making changes to any code in the routes submodule, consult
routes/__init__.pydocstring!
- when making changes to any code in the routes submodule, consult
- schemata: all database schemata, ie, ORM classes, are declared here. There is autodiscovery mechanism, do not make submodules here; only declare schemata in flat
schemata/*.pyfiles, one file per domain (inluding subdomains if any).- do not declare any functions in these files, only the ORM classes themselves
- the user.py represents a domain and a standalone database; the job.py is not a domain but represents a sqlite database
job.dbused by all other domains, and declares only engine/Base, no tables itself. The job.db-residing domains import the Base from jobs.py -- there must be a single Base to handle Foreign Keys correctly.
- domain: the domain entities, related service functions, database helpers, domain dataclasses, et cetera. Most of the business logic lives here. Consult each domain's docstring in
__init__.pyto understand its role. When making any change to a code in a domain, consult the docstring to see if you need to make a change in the docstring itself.- within domain submodules there is often a
db.pywhich contains helper functions to operate on top of ORMs from schemata. This module is always expected to handle automated version increments when mutating versioned entities, and to enforce authorization.- no ORM object may cross the
db.pymodule boundary -- convert query results into a plain dataclass (e.g.RunRecord) before returning. Sessions, active result objects, and lazy-loaded attributes must never be handed to callers; materialize everything the caller needs before the session closes. The ORM objects misbehave when crossing thread boundaries -- which is bound to happen to return values of functions indb.pygiven we mix sync and async code and thread pools.
- no ORM object may cross the
- domains often contain an
exceptions.pyfile -- we prefer domain-specific exceptions over ValueError everywhere. These exceptions should be translated into corresponding HTTP status codes in the route modules.
- within domain submodules there is often a
- utility: code that can be utilized across domains, that is, helper functions operating primarily on standard library constructs
- an exception is utility/config.py, which is containing a lot of domain-specific code. We chose to place it in
utilityto have all config centralized and available to the whole application. - common cross-domain utility exceptions (e.g.
ExecutionManagerErrorand its subclasses,NoFreePortsException) have a general handling in theutility/fastapimodule, wired once inentrypoint/app.py-- rather than being handled ad hoc in every route that might raise them.
- an exception is utility/config.py, which is containing a lot of domain-specific code. We chose to place it in
- entrypoint: code related to bootstrapping the FastAPI backend itself, including self-checks, config management, logging setup, et cetera.
- there are utility-like functions here as well -- when deciding whether to add here or to top-level utility module, consider whether its entrypoint-only or of plausible usage to domains as well
Make sure you don't break importing hierarchies: utility < schemata < domain < routes < entrypoint. There are additional rules for hierarchy within domains -- when you change imports in a particular domain, consult its docstring to understand if that is allowed. In particular, the hierarchy within domains must stay acyclic. If the design seems to require a coupling between two domains, consult the "event dispatcher" in the Concurrency section below. If such situation appears, it still must be noted in the domain's top docstring as a loose/indirect domain coupling.
This application is deployed at multiple machines owned by users, over which we have no control. Changes you make need to preserve compatibility:
- when adding new fields to config.py, make sure they contain defaults -- we need to be backwards compatible wrt users configs. Do not change existing fields -- there are currently no means for migrations.
- there is currently no mechanism for handling migrations -- do not change existing classes in the schemata module. You can add new classes
- when changing anything in the
routessubmodule, consult its docstring for mandatory guidelines
There is currently async loop which serves all the requests to the FastAPI app, as well as multiple background threads and pools: scheduler thread, plugin thread, artifact pool, event dispatcher, database garbage collector. Particular care must be paid to handling things correctly -- we don't necessarily aim to maximize throughput, but we do not want to block the user-facing async loop for long as we need to keep the backend responsiveness high. Additionally, we want some sort of fairness -- one long operation should not starve other short operations.
- sqlite, the jobs persistence layer, supports only one concurrent writer -- jobs-database access is synchronous and serialized by a
threading.RLockinutility/db.py(dbRetry/dbLock). Everydb.pyhelper across domains is a synchronous, operation-local function that acquires this lock internally; it never submits itself to a pool. Implementing support for concurrent reads should be possible in the future, but do not attepmt to solve it now- async code (routes, services) must never call a jobs
db.pyhelper directly -- submit the whole helper call throughexecution_manager.await_jobs_db()(backed by the single-workerConcurrentPools.JobsDbpool). Synchronous code (background threads, pool workers) calls the helper directly, on its own thread, without touching the event loop - the users database (
domain/auth/db.py) is a separate, independent concern -- it keeps its own async lock/retry helper and remains fully async via aiosqlite. Do not mix jobs-database and users-database locking. This database is physically different sqlite file with independent writes/reads, is never accessed outside of the auth web route, and may be migrated to an external auth provider later - it is preferable that a read-modify-write sequence with business logic in between is split into two separate locked callables (two lock acquisitions), rather than one callable holding the lock across the business logic, to keep the lock occupancy low
- async code (routes, services) must never call a jobs
- multiple state structures are updated via the background threads, but consumed by the async loop -- to achieve synchronization, we rely on immutable data structures from the
pyrsistentpackage. All concurrently accessed state is declared as pyrsistent structure, reads are lock-free, and the lock only needs to provide for atomic swap after updates. When working with shared state, make sure you utilize this pattern -- seeManagerclasses in plugin or artifact domains. - threads and thread pools must not be created ad hoc. Use the central
ExecutionManager(utility/concurrency/manager.py'sexecution_managerinstance) for both named bounded pools (ConcurrentPools) and long-lived threads (ConcurrentThreads) -- register at startup with a stage, submit work with asubmit_*method, and provide a non-blocking component status. Do not instantiateThreadPoolExecutororthreading.Threaddirectly in domain code.- some pre-existing background threads (e.g. scheduler, plugin updater, run background submission) still use ad hoc, pre-manager thread creation rather than the manager -- this is known technical debt, not a pattern to copy for new code.
- for one-to-many, loosely coupled reactions (crossing or reversing the usual dependency hierarchy), use the process-local event dispatcher (
utility/dispatcher.py) instead of a direct call or a new thread -- producers publish an immutableEventand never import consumers; the entrypoint discovers and wires handlers via each domain's optionaldispatchers.py. - for a linear workflow where the next step just needs a different pool/lock, prefer
execution_manager.submit_after(a non-blocking continuation) over the event bus -- the event bus is not a substitute for an ordinary same-direction function call.
- sometimes a request to the FastAPI app triggers a possibly long operation, which we offload to a named
ExecutionManagerpool -- for example, submitting a Run domain entity. It is imperative that all such operations are fully wrapped in a try-catch block, and any error manifests by updating the database state with the respective error message.