-
-
Notifications
You must be signed in to change notification settings - Fork 589
S3 File Processor
This guide shows you how to use xyOps and the S3 Transfer Marketplace Plugin to automatically pick up and process files from an S3 bucket. The workflow quietly polls for files, downloads the oldest one, removes it from S3, and launches a visible job to process it.
When there are no files waiting, the polling jobs are invisible and automatically delete themselves. When files are found, the collection jobs are retained, and the processing job runs normally, so you can follow its progress and examine the results.
This is a variation on Remote File Processor: Method 2, Silent Workflow. Here, a single S3 Transfer job handles the pickup, and you can process one file at a time or collect a batch on each poll.
The S3 Transfer Plugin's Download Files tool has three options that work especially well together:
| Parameter | Value | Purpose |
|---|---|---|
| Sort Files | Oldest First (oldest) |
Select the oldest matching files first. |
| Maximum Files | 1 |
Pick up one file per poll. |
| Delete Files | Checked | Remove each source object after it is successfully downloaded. |
Keep Attach Files enabled, and the downloaded files become job outputs that xyOps can pass to another job. A Decision Controller checks the output file count, then a Run Event action launches the processor only when files were collected.
Here is the flow we will build:
flowchart LR
manual[Manual Run] --> pickup[Download Files]
interval[Interval Trigger] --> pickup
pickup -->|On Success| decision{"count(files) > 0"}
decision --> action[Run Event Action]
action -.-> processor[Visible Processing Job]
The processing job runs outside the polling workflow. The Quiet modifier will apply to scheduled polling runs and their pickup sub-jobs. It will run all jobs silently unless files are attached for processing.
Note
Oldest First sorts by the S3 objects' last-modified timestamps. This gives you approximate FIFO pickup, but equal timestamps and overwriting existing objects can affect the order. Use unique object names for incoming files.
You will need:
- xyOps v1.0.10 or newer, for the silent workflow features.
- The S3 Transfer Plugin, installed from the Marketplace in the xyOps sidebar.
- A target server with Node.js and
npxavailable, and network access to S3. - An S3 bucket and credentials that allow listing, downloading, and deleting the incoming objects.
Create a Secret Vault, assign the S3 Transfer Plugin to it, and add these environment variables:
AWS_ACCESS_KEY_IDAWS_SECRET_ACCESS_KEY
For this example, we will use a bucket named my-file-inbox in us-east-1, with incoming files placed directly under the incoming/ prefix. Substitute your own bucket, region, and prefix throughout the guide.
Start by creating a normal event named Process S3 File. This is the event that will receive the downloaded files and perform the actual work.
For an initial test, select the built-in Test Plugin. Choose a target server and leave the event's Manual Run trigger enabled. The Run Event action needs an enabled Manual Run trigger on its destination event.
Add these limits directly to the processing event:
-
Max Jobs Limit:
1 -
Max Queue Limit: A non-zero capacity appropriate for your workload, such as
100.
This allows one processing job to run at a time, while additional pickups wait in the processing queue. Save the event so we can select it in the workflow's Run Event action.
Later, replace the Test Plugin with your real processor. xyOps passes the files into the new job as inputs, and xySat downloads them into that job's temporary directory before the Plugin runs. See Input Files for details.
Create a Workflow named Poll S3 Inbox, starting with a Manual Run trigger node. We will add scheduling and quiet mode after testing the file pickup and handoff.
Add a Job Node, select the S3 Transfer Plugin, and choose the Download Files tool. Give the node a title such as Pick Up S3 File, and select one target server or a server group from which xyOps can choose a single server.
Configure the Plugin parameters like this:
| Parameter | Value |
|---|---|
| Region ID | us-east-1 |
| Bucket Name | my-file-inbox |
| Custom API Endpoint | Leave blank for AWS S3. |
| Remote Path | incoming/ |
| Filename Pattern | Leave blank for all files, or use a pattern such as *.csv. |
| Local Path | Leave blank to use the job's temporary directory. |
| Decompress Files (Gunzip) | Unchecked unless required. |
| Delete Files | Checked |
| Attach Files | Checked |
| Maximum Files | 1 |
| Sort Files | Oldest First (oldest) |
Wire the Manual Run trigger to this node.
Remote Path is an S3 key prefix. Include the trailing slash for a folder-like prefix such as incoming/. The filename pattern applies to filenames, so *.csv limits pickup to CSV files under that prefix.
Important
Delete Files removes the S3 source after the download succeeds, before xyOps attaches the output files and before the processing event runs. It does not wait for successful processing. Once the pickup job completes successfully, its retained output files provide a copy for inspection or recovery, but this flow does not automatically put failed work back into S3.
Wire the S3 Transfer job node to a Decision Controller using the On Success connection condition. Set its expression to:
count(files) > 0The expression uses the previous job's output files. Use the count() helper because arrays in xyOps Expressions do not expose a length property.
If the bucket has no matching files, the pickup succeeds with an empty output file list, the decision evaluates to false, and the workflow ends. If files were downloaded and attached, the decision passes control to the next node.
Make sure you count files, rather than data.files: the former contains the attached job outputs that will be passed into the processor. Attach Files must remain enabled for this recipe.
Wire the Decision Controller to a Run Event action and select Process S3 File as the destination event.
The downloaded files are passed automatically into the new job as inputs. No file paths need to be entered in the action configuration.
Using an action here is the key to making file processing visible. An ordinary Event or Job node inside the polling workflow would inherit its quiet setting and remain hidden. The Run Event action launches a separate job outside that workflow, so it runs visibly and keeps its normal job history.
Set the Max Jobs Limit on the polling workflow itself to 1. This prevents two runs of this workflow from listing and downloading the same S3 objects at the same time. Configure a small Max Queue Limit, such as 1, to allow one pending poll to wait.
Use just one polling workflow for this bucket and prefix, and keep manual tests from overlapping with another consumer of the same objects. Download-and-delete is a sequence of operations, so separate workflows or external consumers could still pick up the same file concurrently.
The polling and processing limits have separate scopes. The workflow finishes after handing off the files; it does not wait for the processing job to finish. That is why the processing event also has its own Max Jobs Limit and queue capacity.
Note
When a queue fills, additional jobs can abort instead of waiting. This matters especially for processing jobs, because their source files have already been consumed from S3. Size the processing queue for your expected backlog, and choose a polling interval that your pickup and processing jobs can keep up with.
Keep the workflow unscheduled while you test. Use disposable files under your chosen prefix, since Delete Files is enabled.
- Run the workflow manually with no matching files in the prefix. The pickup should succeed, and no processing job should launch.
- Upload two test files at different times, then run the workflow again. It should download and delete only the older file.
- Examine the pickup job's output files and the Test Plugin job's inputs. The same file should appear in both.
- Run the workflow once more to collect the remaining file.
- Run it again with the prefix empty, and confirm that no additional processing job launches.
Manual runs remain visible, which makes it easy to inspect the workflow and its results before enabling quiet mode.
Once the manual tests work, add an Interval trigger to the workflow and wire it to the same pickup node. For example, poll every 60 seconds, or choose a longer interval to suit your workload. Keep the Manual Run trigger for testing.
Add a Quiet Modifier node with both options enabled:
- Invisible: Hide upcoming, queued, and running polling jobs from the UI.
- Ephemeral: Automatically delete completed polling jobs when they have no output files and did not fail.
The Quiet node is a schedule modifier and does not need any wires. It applies when the workflow launches from its scheduler trigger, and does not apply to manual runs. The scheduled workflow passes these settings down to its pickup sub-job.
Here is what you should expect:
| Result | Behavior |
|---|---|
| No matching files | The workflow and pickup job run invisibly, then delete themselves on completion. No processing job launches. |
| Files collected | The pickup job and parent workflow are retained because they have output files, and become visible upon completion. The processing job launches as a normal visible job. |
| Pickup fails | The failed pickup and parent workflow are retained for inspection. The On Success connection prevents the processor from launching. |
Producing output files disables ephemeral mode upon completion. This preserves the job that owns the attachments, and gives you a way to trace processing back to the S3 pickup. A job failure also disables ephemeral mode, so errors remain available for troubleshooting.
To collect a batch, increase Maximum Files on the Download Files node. For example, 10 picks up the ten oldest matching objects on each poll, or fewer if less are available. A value of 0 means no limit.
The decision expression stays the same. The Run Event action launches one processing job with all of the collected files as inputs.
Your processor can handle the whole batch, or it can be a separate workflow that uses a Split Controller with the expression files and a Batch Size of 1. Connect its Manual Run trigger directly to the Split Controller so the input files reach it, then attach job and queue limits to the node that performs the per-file work.
Keep the S3 polling workflow limited to one concurrent run even when collecting batches. If processing order matters, keep processing serial as well. Parallel processing can finish files in a different order from their pickup order.