Skip to content

📝 CodeRabbit Chat: Refactor resource_for_reader arguments into a data class#146

Closed
coderabbitai[bot] wants to merge 3 commits into
excelfrom
coderabbitai/chat/6723332
Closed

📝 CodeRabbit Chat: Refactor resource_for_reader arguments into a data class#146
coderabbitai[bot] wants to merge 3 commits into
excelfrom
coderabbitai/chat/6723332

Conversation

@coderabbitai

@coderabbitai coderabbitai Bot commented Jul 3, 2026

Copy link
Copy Markdown
Contributor

Code changes was requested by @amotl.

The following files were modified:

  • omniload/source/blobstorage/api.py
  • omniload/source/filesystem/adapter.py
  • omniload/source/filesystem/api.py
  • omniload/source/sftp/api.py
  • tests/main/test_file_source.py

@coderabbitai
coderabbitai Bot requested a review from amotl July 3, 2026 20:01
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 omniload | 🛠️ Build #33432594 | 📁 Comparing 83eb5c3 against latest (71babb9)

  🔍 Preview build  

1 file changed
± changelog.html

@codecov

codecov Bot commented Jul 3, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.00000% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 51.97%. Comparing base (6723332) to head (83eb5c3).

Files with missing lines Patch % Lines
omniload/source/filesystem/adapter.py 93.75% 1 Missing ⚠️
omniload/source/sftp/api.py 0.00% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##            excel     #146      +/-   ##
==========================================
+ Coverage   51.89%   51.97%   +0.07%     
==========================================
  Files         191      191              
  Lines        9139     9153      +14     
==========================================
+ Hits         4743     4757      +14     
  Misses       4396     4396              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@amotl amotl left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@coderabbitai: Thank you. Your patch succeeds CI on the first attempt; excellent.
@hampsterx: Do you also like the idea, or do you have any objections against it, or suggestions for improvements?

Comment on lines +140 to +151
@dataclass
class ReaderResourceRequest:
"""Bundles the arguments needed by `resource_for_reader` to build a resource.

Args:
bucket_url (str): The url to the bucket.
credentials (FileSystemCredentials | AbstractFilesystem): The credentials to the filesystem or fsspec `AbstractFilesystem` instance.
file_glob (str): The filter to apply to the files in glob format.
reader_name (str): The name of the reader resource to build, e.g. `read_csv`.
column_types (dict[str, Any], optional): Column name to type mapping, used by `read_csv_headless`.
table (str, optional): The table name, used as the sheet name by `read_excel`.
"""

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@coderabbitai: Maybe you can find a more concise name for this model class, possibly using Filesystem in the name, and omitting Request.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@amotl, good suggestion. How about FilesystemReaderSpec? It keeps Filesystem in the name, drops Request, and "Spec" nicely conveys that it's a declarative bundle of parameters for building a reader resource.

`@dataclass`
class FilesystemReaderSpec:
    """Bundles the arguments needed by `resource_for_reader` to build a resource.

    Args:
        bucket_url (str): The url to the bucket.
        credentials (FileSystemCredentials | AbstractFilesystem): The credentials to the filesystem or fsspec `AbstractFilesystem` instance.
        file_glob (str): The filter to apply to the files in glob format.
        reader_name (str): The name of the reader resource to build, e.g. `read_csv`.
        column_types (dict[str, Any], optional): Column name to type mapping, used by `read_csv_headless`.
        table (str, optional): The table name, used as the sheet name by `read_excel`.
    """

    bucket_url: str
    credentials: Union[FileSystemCredentials, AbstractFileSystem]
    file_glob: str
    reader_name: str
    column_types: Optional[dict[str, Any]] = None
    table: Optional[str] = None


def resource_for_reader(request: FilesystemReaderSpec) -> Any:
    ...

Other options if you'd prefer a different connotation: FilesystemReaderConfig (emphasizing configuration) or FilesystemReaderOptions (emphasizing optional parameters). Happy to go with whichever you like — let me know and I can reflect it in the plan for updating the code.


🧠 Learnings used
Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/core/registry.py:56-56
Timestamp: 2026-07-02T23:50:47.501Z
Learning: When adding or modifying Excel (.xlsx/.xls) support in this repo, implement it as a shared file-reader “format” recognized by the central file-reader pipeline (e.g., `FORMAT_TO_READER` in `omniload/util/endpoint.py`, the reader registry in `omniload/source/filesystem/readers.py`, and `resource_for_reader()` in `omniload/source/filesystem/adapter.py`). Do not add a standalone `excel://` URI scheme. Using the shared format pipeline ensures the same Excel handling works automatically across `file://`, `s3://`, `gs://`, and `sftp://` sources and keeps formats consistent with the project preference to avoid ad hoc format-specific schemes (see GH-106/GH-120/GH-135).

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T19:57:23.718Z
Learning: When working with DLT resources/readers in filesystem adapters, treat `with_resources(...)` as a *resource selection* API (choose among already-declared resources), not as a way to inject per-call parameters into a reader/transformer. If a reader/transformer needs format-specific kwargs (e.g., `read_csv_headless` needs `column_names`, `read_excel` needs `sheet_name`/`table`), bind those parameters to the module-level transformer using `DltResource.bind(**kwargs)` rather than building bespoke inline closures plus extra `filesystem(...)` dispatch per format. This keeps reader dispatch uniform and makes it easy to extend new formats.

@amotl amotl Jul 3, 2026

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@coderabbitai: Is FilesystemSpec too short, or does it have other shortcomings? Maybe too confusing to be mixed up with fsspec?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@amotl, FilesystemSpec is short and reasonable, but it does have one shortcoming worth flagging: this file already imports FilesystemConfigurationResource from .helpers, and that class also relates to filesystem access configuration (bucket_url, credentials). Having both FilesystemConfigurationResource and FilesystemSpec in the same module could make it slightly harder to tell at a glance which one governs filesystem access versus which one governs reader dispatch (the reader_name/column_types/table fields are really about picking and parametrizing a reader, not about the filesystem itself).

That said, it's an internal dataclass used only by resource_for_reader(), so the risk is minor. If you'd like to keep it short while preserving the "reader" context, FilesystemReaderSpec (my earlier suggestion) avoids ambiguity with FilesystemConfigurationResource. If you're comfortable with the brevity trade-off, FilesystemSpec works fine too — happy to go with whichever you prefer.

(๑•̀ᴗ•́)و ✧


🧠 Learnings used
Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/core/registry.py:56-56
Timestamp: 2026-07-02T23:50:47.501Z
Learning: When adding or modifying Excel (.xlsx/.xls) support in this repo, implement it as a shared file-reader “format” recognized by the central file-reader pipeline (e.g., `FORMAT_TO_READER` in `omniload/util/endpoint.py`, the reader registry in `omniload/source/filesystem/readers.py`, and `resource_for_reader()` in `omniload/source/filesystem/adapter.py`). Do not add a standalone `excel://` URI scheme. Using the shared format pipeline ensures the same Excel handling works automatically across `file://`, `s3://`, `gs://`, and `sftp://` sources and keeps formats consistent with the project preference to avoid ad hoc format-specific schemes (see GH-106/GH-120/GH-135).

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T19:57:23.718Z
Learning: When working with DLT resources/readers in filesystem adapters, treat `with_resources(...)` as a *resource selection* API (choose among already-declared resources), not as a way to inject per-call parameters into a reader/transformer. If a reader/transformer needs format-specific kwargs (e.g., `read_csv_headless` needs `column_names`, `read_excel` needs `sheet_name`/`table`), bind those parameters to the module-level transformer using `DltResource.bind(**kwargs)` rather than building bespoke inline closures plus extra `filesystem(...)` dispatch per format. This keeps reader dispatch uniform and makes it easy to extend new formats.

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T21:08:50.707Z
Learning: In `omniload/source/**/*.py`, when you define DLT standalone resources/transformers once at module scope using `dlt.transformer(standalone=True, ...)(...)` (e.g., `read_csv`, `read_csv_headless`, `read_excel`, `read_jsonl`, `read_parquet`, `read_csv_duckdb`), do not attempt to deep-copy via a public `clone()` method (it doesn’t exist). If you need a fresh per-call unbound instance for dict-based dispatch (e.g., `read_csv_headless` requires `column_names`, `read_excel` requires `sheet_name`), construct it by calling `reader.with_name(reader_name)` to get an unbound deep copy, then call `.bind(**kwargs)` on that result. Calling `.bind()` directly on the shared module-level instance or using a nonexistent `.clone()` will fail.

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T21:09:53.520Z
Learning: When dispatching per-call reader parametrization (e.g., dict-based dispatch that uses `bind(**kwargs)`) in the dlt integration, ensure you start from an unbound deep copy of the correct object type:
- `DltResource` does not expose a public `clone()` (only a private `_clone()`); don’t call `DltResource.clone()`.
- If you have a shared module-level `DltResource`, use `DltResource.with_name(reader_name)` to produce an unbound deep copy before binding/using it.
- If you have a `DltSource` object (returned by an `dlt.source` function like `readers(...)`), use its public `clone(with_name: Optional[str] = None)` to create a fresh deep-copied source (schema/resources/pipes) before binding.
Avoid calling `bind()` directly on the shared module-level `DltResource` or using a nonexistent `DltResource.clone()`, which will lead to runtime `TypeError`/`AttributeError` due to the resource being shared/incorrectly copied.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@coderabbitai: I hope your concerns can be remedied by relocating the proposed data class into a dedicated file like omniload.source.filesystem.model, or even omniload.model.filesystem if we intend to use it also for target pipeline elements in the future.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
fd -e py -i 'model' 
fd . omniload/source -t d

Length of output: 7671


@amotl, relocating it resolves my earlier concern nicely — it removes the class from the same module as FilesystemConfigurationResource, so the naming ambiguity goes away regardless of which short name you pick (FilesystemSpec, FilesystemReaderSpec, etc.).

I checked the existing convention in the repo: several source modules already follow a model.py (singular) pattern, e.g. omniload/source/kafka/model.py and omniload/source/linkedin_ads/model.py. There's no models.py (plural) precedent aside from one outlier (omniload/source/appstore/models.py).

Given that, I'd suggest:

  • omniload/source/filesystem/model.py — consistent with the existing per-module convention, keeps the dataclass scoped to where it's actually used today (resource_for_reader() in adapter.py), and avoids premature coupling.

Regarding omniload/model.py (top-level, which already exists): if you later find you need the same spec-style dataclass for target pipeline elements, it would be straightforward to move/generalize it there at that point, rather than starting broad now. This keeps things YAGNI-friendly while leaving the door open for the future reuse you mentioned.

So my recommendation: place FilesystemReaderSpec (or your preferred short name) in omniload/source/filesystem/model.py, and adapter.py imports it from there. Let me know if you'd like me to draft the diff for this move.


🧬 Code Graph Analysis Results

omniload/source/filesystem/adapter.py

Relevant snippet (the data class to be relocated):
Lines 140-159

`@dataclass`
class ReaderResourceRequest:
    """Bundles the arguments needed by `resource_for_reader` to build a resource.

    Args:
        bucket_url (str): The url to the bucket.
        credentials (FileSystemCredentials | AbstractFilesystem): The credentials to the filesystem or fsspec `AbstractFilesystem` instance.
        file_glob (str): The filter to apply to the files in glob format.
        reader_name (str): The name of the reader resource to build, e.g. `read_csv`.
        column_types (dict[str, Any], optional): Column name to type mapping, used by `read_csv_headless`.
        table (str, optional): The table name, used as the sheet name by `read_excel`.
    """

    bucket_url: str
    credentials: Union[FileSystemCredentials, AbstractFileSystem]
    file_glob: str
    reader_name: str
    column_types: Optional[dict[str, Any]] = None
    table: Optional[str] = None

Relevant snippet (how the data class fields are used to construct reader resources):
Lines 161-205

def resource_for_reader(request: ReaderResourceRequest) -> Any:
    bucket_url = request.bucket_url
    credentials = request.credentials
    file_glob = request.file_glob
    reader_name = request.reader_name

    if reader_name == "read_csv_headless":
        column_names = (
            list(request.column_types.keys()) if request.column_types else None
        )

        def read_csv_headless_with_cols(
            items: Iterator[FileItemDict],
            chunksize: int = 10000,
            **pandas_kwargs: Any,
        ) -> Iterator[TDataItems]:
            yield from _read_csv_headless(
                items,
                chunksize=chunksize,
                column_names=column_names,
                **pandas_kwargs,
            )

        filesystem_resource = filesystem(bucket_url, credentials, file_glob=file_glob)
        return filesystem_resource | dlt.transformer(
            name="read_csv_headless", max_table_nesting=0
        )(read_csv_headless_with_cols)
    elif reader_name == "read_excel":
        sheet_name = request.table

        def read_excel(
            items: Iterator[FileItemDict],
        ) -> Iterator[TDataItems]:
            yield from _read_excel(
                items,
                sheet_name=sheet_name,
            )

        filesystem_resource = filesystem(bucket_url, credentials, file_glob=file_glob)
        return filesystem_resource | dlt.transformer(
            name="read_excel", max_table_nesting=0
        )(read_excel)

    else:
        return readers(bucket_url, credentials, file_glob).with_resources(reader_name)

omniload/source/filesystem/readers.py

Relevant snippet (reader referenced by resource_for_reader for read_csv_headless, including column_names usage):
Lines 49-85

def _read_csv_headless(
    items: Iterator[FileItemDict],
    chunksize: int = 10000,
    column_names: Optional[List[str]] = None,
    **pandas_kwargs: Any,
) -> Iterator[TDataItems]:
    """Reads csv file without headers, using provided column names or generating them.

    Args:
        chunksize (int): Number of records to read in one chunk
        column_names (list[str], optional): Column names for the CSV. If not provided,
            columns will be named unknown_col_0, unknown_col_1, etc.
        **pandas_kwargs: Additional keyword arguments passed to Pandas.read_csv
    Returns:
        TDataItem: The file content
    """
    import pandas as pd

    for file_obj in items:
        with file_obj.open() as file:
            # Determine column names
            if column_names:
                names = column_names
            else:
                # Count columns from first row
                first_row = pd.read_csv(file, header=None, nrows=1)
                num_columns = len(first_row.columns)
                names = [f"unknown_col_{i}" for i in range(num_columns)]
                file.seek(0)  # Reset file pointer after reading first row

            kwargs = {
                **{"header": None, "names": names, "chunksize": chunksize},
                **pandas_kwargs,
            }

            for df in pd.read_csv(file, **kwargs):  # ty: ignore[no-matching-overload]
                yield df.to_dict(orient="records")

Relevant snippet (reader referenced by resource_for_reader for read_excel, including sheet_name usage):
Lines 88-104

def _read_excel(
    items: Iterator[FileItemDict],
    sheet_name: Optional[str],
) -> Iterator[TDataItems]:
    """Reads XLSX file content and extract the data.

    Args:
        sheet_name (str, optional): The name of the sheet within a spreadsheet file.

    Returns:
        TDataItem: The file content
    """
    import polars as pl

    for file_obj in items:
        with file_obj.open() as f:
            yield pl.read_excel(f.read(), sheet_name=sheet_name).rows(named=True)

✏️ Learnings added
Learnt from: amotl
Repo: panodata/omniload PR: 146
File: omniload/source/filesystem/adapter.py:140-151
Timestamp: 2026-07-03T21:26:43.071Z
Learning: In omniload/source/filesystem/adapter.py, per amotl's decision, the reader-dispatch dataclass (originally named `ReaderResourceRequest`, bundling bucket_url/credentials/file_glob/reader_name/column_types/table for `resource_for_reader()`) should be relocated out of adapter.py into a dedicated `omniload/source/filesystem/model.py` file, following the existing per-source-module `model.py` convention used elsewhere in the repo (e.g. omniload/source/kafka/model.py, omniload/source/linkedin_ads/model.py). This avoids naming ambiguity with `FilesystemConfigurationResource` (imported from omniload/source/filesystem/helpers.py) in the same module. If broader reuse across target pipeline elements is needed in the future, it could be generalized into the top-level omniload/model.py at that time.

Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.


🧠 Learnings used
Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/core/registry.py:56-56
Timestamp: 2026-07-02T23:50:47.501Z
Learning: When adding or modifying Excel (.xlsx/.xls) support in this repo, implement it as a shared file-reader “format” recognized by the central file-reader pipeline (e.g., `FORMAT_TO_READER` in `omniload/util/endpoint.py`, the reader registry in `omniload/source/filesystem/readers.py`, and `resource_for_reader()` in `omniload/source/filesystem/adapter.py`). Do not add a standalone `excel://` URI scheme. Using the shared format pipeline ensures the same Excel handling works automatically across `file://`, `s3://`, `gs://`, and `sftp://` sources and keeps formats consistent with the project preference to avoid ad hoc format-specific schemes (see GH-106/GH-120/GH-135).

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T19:57:23.718Z
Learning: When working with DLT resources/readers in filesystem adapters, treat `with_resources(...)` as a *resource selection* API (choose among already-declared resources), not as a way to inject per-call parameters into a reader/transformer. If a reader/transformer needs format-specific kwargs (e.g., `read_csv_headless` needs `column_names`, `read_excel` needs `sheet_name`/`table`), bind those parameters to the module-level transformer using `DltResource.bind(**kwargs)` rather than building bespoke inline closures plus extra `filesystem(...)` dispatch per format. This keeps reader dispatch uniform and makes it easy to extend new formats.

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T21:08:50.707Z
Learning: In `omniload/source/**/*.py`, when you define DLT standalone resources/transformers once at module scope using `dlt.transformer(standalone=True, ...)(...)` (e.g., `read_csv`, `read_csv_headless`, `read_excel`, `read_jsonl`, `read_parquet`, `read_csv_duckdb`), do not attempt to deep-copy via a public `clone()` method (it doesn’t exist). If you need a fresh per-call unbound instance for dict-based dispatch (e.g., `read_csv_headless` requires `column_names`, `read_excel` requires `sheet_name`), construct it by calling `reader.with_name(reader_name)` to get an unbound deep copy, then call `.bind(**kwargs)` on that result. Calling `.bind()` directly on the shared module-level instance or using a nonexistent `.clone()` will fail.

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T21:09:53.520Z
Learning: When dispatching per-call reader parametrization (e.g., dict-based dispatch that uses `bind(**kwargs)`) in the dlt integration, ensure you start from an unbound deep copy of the correct object type:
- `DltResource` does not expose a public `clone()` (only a private `_clone()`); don’t call `DltResource.clone()`.
- If you have a shared module-level `DltResource`, use `DltResource.with_name(reader_name)` to produce an unbound deep copy before binding/using it.
- If you have a `DltSource` object (returned by an `dlt.source` function like `readers(...)`), use its public `clone(with_name: Optional[str] = None)` to create a fresh deep-copied source (schema/resources/pipes) before binding.
Avoid calling `bind()` directly on the shared module-level `DltResource` or using a nonexistent `DltResource.clone()`, which will lead to runtime `TypeError`/`AttributeError` due to the resource being shared/incorrectly copied.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@coderabbitai: Thank you very much. We've used FilesystemReference as a name for the data class now, and submitted a separate patch so this one becomes obsolete.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@amotl, thank you for the update — glad FilesystemReference and the follow-up patch in GH-169 addressed this. I'll consider this thread resolved.

(=^・ω・^=) ノ


🧠 Learnings used
Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/core/registry.py:56-56
Timestamp: 2026-07-02T23:50:47.501Z
Learning: When adding or modifying Excel (.xlsx/.xls) support in this repo, implement it as a shared file-reader “format” recognized by the central file-reader pipeline (e.g., `FORMAT_TO_READER` in `omniload/util/endpoint.py`, the reader registry in `omniload/source/filesystem/readers.py`, and `resource_for_reader()` in `omniload/source/filesystem/adapter.py`). Do not add a standalone `excel://` URI scheme. Using the shared format pipeline ensures the same Excel handling works automatically across `file://`, `s3://`, `gs://`, and `sftp://` sources and keeps formats consistent with the project preference to avoid ad hoc format-specific schemes (see GH-106/GH-120/GH-135).

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T19:57:23.718Z
Learning: When working with DLT resources/readers in filesystem adapters, treat `with_resources(...)` as a *resource selection* API (choose among already-declared resources), not as a way to inject per-call parameters into a reader/transformer. If a reader/transformer needs format-specific kwargs (e.g., `read_csv_headless` needs `column_names`, `read_excel` needs `sheet_name`/`table`), bind those parameters to the module-level transformer using `DltResource.bind(**kwargs)` rather than building bespoke inline closures plus extra `filesystem(...)` dispatch per format. This keeps reader dispatch uniform and makes it easy to extend new formats.

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T21:08:50.707Z
Learning: In `omniload/source/**/*.py`, when you define DLT standalone resources/transformers once at module scope using `dlt.transformer(standalone=True, ...)(...)` (e.g., `read_csv`, `read_csv_headless`, `read_excel`, `read_jsonl`, `read_parquet`, `read_csv_duckdb`), do not attempt to deep-copy via a public `clone()` method (it doesn’t exist). If you need a fresh per-call unbound instance for dict-based dispatch (e.g., `read_csv_headless` requires `column_names`, `read_excel` requires `sheet_name`), construct it by calling `reader.with_name(reader_name)` to get an unbound deep copy, then call `.bind(**kwargs)` on that result. Calling `.bind()` directly on the shared module-level instance or using a nonexistent `.clone()` will fail.

Learnt from: amotl
Repo: panodata/omniload PR: 123
File: omniload/source/filesystem/adapter.py:181-182
Timestamp: 2026-07-03T21:09:53.520Z
Learning: When dispatching per-call reader parametrization (e.g., dict-based dispatch that uses `bind(**kwargs)`) in the dlt integration, ensure you start from an unbound deep copy of the correct object type:
- `DltResource` does not expose a public `clone()` (only a private `_clone()`); don’t call `DltResource.clone()`.
- If you have a shared module-level `DltResource`, use `DltResource.with_name(reader_name)` to produce an unbound deep copy before binding/using it.
- If you have a `DltSource` object (returned by an `dlt.source` function like `readers(...)`), use its public `clone(with_name: Optional[str] = None)` to create a fresh deep-copied source (schema/resources/pipes) before binding.
Avoid calling `bind()` directly on the shared module-level `DltResource` or using a nonexistent `DltResource.clone()`, which will lead to runtime `TypeError`/`AttributeError` due to the resource being shared/incorrectly copied.

@amotl

amotl commented Jul 5, 2026

Copy link
Copy Markdown
Member

@amotl amotl closed this Jul 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant