Skip to content

utils

PlantDB Utility FunctionsLink

A small collection of helper utilities used throughout the PlantDB project to simplify common filesystem-related tasks. These functions handle safe retrieval of database resources, verify path containment, and inspect the contents of ZIP archives, reducing boiler-plate and ensuring consistent error handling across the code base.

Key FeaturesLink

  • resource_file: Retrieves a File object from the PlantDB database, translating common error conditions into clear JSON-style messages and HTTP status codes.
  • is_within_directory: Checks whether a target path lies inside a given directory, preventing directory-traversal vulnerabilities.
  • is_directory_in_archive: Determines if a specific top-level directory exists inside a ZIP archive, useful for validating package structures before extraction.

Usage ExamplesLink

from plantdb.server.api.utils import resource_file, is_within_directory, is_directory_in_archive

Retrieve a file from the database (returns a File object or an error dict)Link

result = resource_file(db, "scan123", "segmentation")

Verify a path is under a base directoryLink

is_within_directory("/data/plantdb", "/data/plantdb/scans/scan123") True

Check for a directory named 'images' inside a zip fileLink

is_directory_in_archive("dataset.zip", "images") True

is_directory_in_archive Link

is_directory_in_archive(archive_path, target_dir)

Check if a specific directory exists within an archive file.

This function checks whether a given directory is present at the top level of a ZIP archive.

Parameters:

Name Type Description Default

archive_path Link

str or Path

The path to the ZIP archive file.

required

target_dir Link

str

The name of the target directory to check for within the archive.

required

Returns:

Type Description
bool

True if the target directory exists at the top level of the archive, False otherwise.

Source code in plantdb/server/api/utils.py
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
def is_directory_in_archive(archive_path, target_dir):
    """Check if a specific directory exists within an archive file.

    This function checks whether a given directory is present at the top level of a ZIP archive.

    Parameters
    ----------
    archive_path : str or pathlib.Path
        The path to the ZIP archive file.
    target_dir : str
        The name of the target directory to check for within the archive.

    Returns
    -------
    bool
        True if the target directory exists at the top level of the archive, False otherwise.
    """
    with ZipFile(archive_path, 'r') as zip_ref:
        # List all members in the zip file
        top_level_members = [name for name in zip_ref.namelist() if '/' not in name]
        # Check if the target directory is among them
        return f"{target_dir}/" in top_level_members or target_dir in top_level_members

is_within_directory Link

is_within_directory(directory, target)

Check if a target path is within a directory.

This function determines if the absolute path of the target is located within the absolute path of the directory. It uses os.path.commonpath to perform the comparison.

Parameters:

Name Type Description Default

directory Link

str or Path

The path to the directory to check against.

required

target Link

str or Path

The path to the target to check if it resides within the directory.

required

Returns:

Type Description
bool

True if the target path is within the directory, False otherwise.

Source code in plantdb/server/api/utils.py
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
def is_within_directory(directory, target):
    """Check if a target path is within a directory.

    This function determines if the absolute path of the target is located
    within the absolute path of the directory. It uses `os.path.commonpath`
    to perform the comparison.

    Parameters
    ----------
    directory : str or pathlib.Path
        The path to the directory to check against.
    target : str or pathlib.Path
        The path to the target to check if it resides within the directory.

    Returns
    -------
    bool
        ``True`` if the target path is within the directory, ``False`` otherwise.
    """
    abs_directory = os.path.abspath(directory)
    abs_target = os.path.abspath(target)
    return os.path.commonpath([abs_directory]) == os.path.commonpath([abs_directory, abs_target])

resource_file Link

resource_file(db, scan_id, task_name, **kwargs)

Retrieve a specific File object from the database.

Parameters:

Name Type Description Default

db Link

FSDB

The database instance.

required

scan_id Link

str

Identifier of the scan containing the requested file.

required

task_name Link

str

Name of the task (fileset) and file to fetch.

required

**kwargs Link

Additional arguments passed to FSDB.get_scan (e.g. JWT token).

{}

Returns:

Type Description
tuple

Either (file, 200) on success or (error_dict, status_code) on failure. The error dictionary always uses the key "error" for consistency.

Source code in plantdb/server/api/utils.py
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
def resource_file(db, scan_id, task_name, **kwargs):
    """Retrieve a specific ``File`` object from the database.

    Parameters
    ----------
    db : plantdb.commons.fsdb.core.FSDB
        The database instance.
    scan_id : str
        Identifier of the scan containing the requested file.
    task_name : str
        Name of the task (fileset) and file to fetch.
    **kwargs
        Additional arguments passed to ``FSDB.get_scan`` (e.g. JWT token).

    Returns
    -------
    tuple
        Either ``(file, 200)`` on success or ``(error_dict, status_code)`` on failure.
        The error dictionary always uses the key ``"error"`` for consistency.
    """

    # Get the corresponding `Scan` instance
    try:
        scan = db.get_scan(scan_id, **kwargs)

    except NoAuthUserError as e:
        return {'error': str(e)}, 401  # HTTP 401 Unauthorized (authentication)
    except ScanNotFoundError:
        return {"error": f"Scan '{scan_id}' not found!"}, 400

    task_fs_map = compute_fileset_matches(scan)
    # Get the corresponding `Fileset` instance
    try:
        fs = scan.get_fileset(task_fs_map[task_name])
    except KeyError:
        return {"error": f"No fileset mapped for task '{task_name}'."}, 404
    except FilesetNotFoundError:
        return {"error": f"Fileset for task '{task_name}' not found."}, 404

    # Get the `File` corresponding to the resource
    try:
        file = fs.get_file(task_name)
    except FileNotFoundError:
        return {"error": f"File '{fs.id}/{task_name}' not found."}, 404
    except Exception as exc:                     # Unexpected internal error
        # Use JSON-serializable payload; Flask will handle conversion.
        return {"error": f"Internal server error: {str(exc)}"}, 500

    # Success: return the File object (Flask-RESTful resources expect the
    # object itself; the caller can decide the HTTP status if required).
    return file