Replies: 2 comments 1 reply
|
I have created a hook that does the job: def remove_duplicates(extracted_entries_list, ledger_entries):
filtered_entries = []
for filename, entries, account, importer in extracted_entries_list:
removed_dups = [entry for entry in entries if not entry.meta.pop(DUPLICATE, False)]
filtered_entries.append((filename, removed_dups, account, importer))
return filtered_entries |
0 replies
|
Transactions (and other ledger entries to which metadata can be attached) marked with the special |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hello!
I would like to discuss, what is the best place to perform deduplication. Here, I mean deduplication in the sense to avoid importing the same entry twice, e.g. when importing overlapping statements.
The default implementation marks entries with meta data if they are heuristically similar.
As I am saving the original line from CSV file in the meta data, I can use that line to compare. Writing a suitable comparator is not a problem.
The duplicate function however can only modify transactions, not delete them AFAIK.
What is the right place to remove transactions that are duplicates? Is there a slot for a post-filtering or do I need to overwrite the extract function?
Best Thanks!
All reactions