Ideas #11
Replies: 4 comments 1 reply
|
I would be interested in being able to use custom filter chain arguments (-vf ...) to ffmpeg toward the goal of using the dnn_classify/dnn_detect/dnn_processing and draw_box filter to apply models for purposes other than upscaling. My particular use case involves using models like YOLO-v3-tiny (or other similar models) from the openvino open model zoo to do object detection/classification and then draw boxes to cover specific detected objects. For me the goal of this is to block animals (but could be any classification supported by the model) from appearing on the videos by overlaying a solid colored box because my lovely dog that watches TV with me loses her mind if she sees a dog or cat on the screen. I can achieve this type of result on videos externally (see attached) but it would be very helpful to have it as a real-time filter. local-file-processed-1785819074497.mp4 |
|
Hi,
Thank you for taking the time to explain. I had looked at the
"camera-style" filters and saw the ffmpeg filter chains listed and
incorrectly assumed that those were also responsible for the upscaling
model application. I also forgot that jellyfin ran its own fork of ffmpeg.
Your approach seems more robust.
I am trying to test the new features you added for v1.8.3.24 but ran into
an issue with the model import. I am using the
kuscheltier/jellyfin-ai-upscaler:docker7-cpu image. I get this error when
trying to install the model through the jellyfin plugin controller:
Install failed: Invalid ONNX model shape: Expected 4D output (N, C, H, W),
got shape [1, None, 4]
I assume this is related to the shape of the model inputs and outputs as
you described:
It is not a plain YOLO head: it takes two inputs (the image plus the
original frame size) and returns three tensors (boxes, per-class scores,
NMS-selected indices), with NMS and the letterbox correction done inside
the graph and boxes coming back as (y_min, x_min, y_max, x_max) in source
coordinates. My decoder fed one input and read output[0] as a single
combined tensor.
If I haven't made some error in the model installation process, perhaps
there needs to be some way to override this check or take some manual input
about the model input/output shapes?
Thanks in advance.
…On Tue, Aug 4, 2026 at 6:59 PM Kuschel-code ***@***.***> wrote:
@schmore <https://github.com/schmore> Follow-up: *v1.8.3.24 does the
real-time part*, which is the half of your request I did not deliver last
time.
I was straight about that gap in my previous reply and then went and
closed it. What v1.8.3.23 shipped was a per-frame endpoint you had to drive
yourself — but you could already produce that result externally. Live
during playback was the actual request.
It turned out to be a small change, because the machinery was already
there: the player captures frames off the <video> element for server-side
upscaling and draws the result onto an overlay. Masking simply had nowhere
to send them.
*How to use it*
1. Import a detection model on the *Models* tab — tiny-yolov3-11.onnx
from the ONNX model zoo is the one you linked, and it is supported.
2. *Settings → Object Masking*: enable it, enter the model id, set
classes (animals by default), pick box or blur, then save and press *Load
Detection Model*.
3. Press play. You can also toggle it mid-session from the player menu
under *Auto → Cover objects*.
*One thing to know before you try it:* while masking runs it *replaces*
real-time upscaling on that stream. Upscaling and detecting is two full
inference passes per frame, and no realistic server keeps that up at
playback rate — so rather than let you discover it as stutter, the stream
does one or the other. Batch upscaling is unaffected.
Padding defaults to 12 px, and that default exists because of your
specific problem: a detector's box hugs the animal, and the ears and tail
sticking out of a tight box will set her off just as well as the whole dog.
Raise it if she still reacts.
*What I would genuinely like to know*, since I cannot test this part
myself — I have no dog and no detection weights on this machine, so
everything here is verified against synthesised models rather than real
ones:
- Does tiny-yolov3-11.onnx load and produce boxes in the right places
on your material?
- What frame rate do you get, and on what hardware?
- Does the default padding actually cover enough for her?
If the boxes land wrong or the rate is unusable, that is worth its own
issue and I will look at it.
—
Reply to this email directly, view it on GitHub
<#11?email_source=notifications&email_token=CKPAU75LE3PUOJIBBKDDVU35IKBGZA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZZGAYDCMJYUZZGKYLTN5XKO3LFNZ2GS33OUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-17900118>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/CKPAU7Y3FGYM7P47TCHXB3L5IKBGZAVCNFSNUABHKJSXA33TNF2G64TZHM4DSMJQGUZTSNBZHNCGS43DOVZXG2LPNY5TQNJWHE4TSOFBOYBA>
.
Triage notifications, keep track of coding agent tasks and review pull
requests on the go with GitHub Mobile for iOS
<https://github.com/notifications/mobile/ios/CKPAU7276EZGZ5FBNLNMNZD5IKBGZA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZZGAYDCMJYUZZGKYLTN5XKO3LFNZ2GS33OUVSXMZLOOSVGM33PORSXEX3JN5ZQ>
and Android
<https://github.com/notifications/mobile/android/CKPAU7Y7VUCKOJP5QWH5SQT5IKBGZA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCNZZGAYDCMJYUZZGKYLTN5XKO3LFNZ2GS33OUVSXMZLOOSXGM33PORSXEX3BNZSHE33JMQ>.
Download it today!
You are receiving this because you were mentioned.Message ID:
<Kuschel-code/JellyfinUpscalerPlugin/repo-discussions/11/comments/17900118
@github.com>
|
|
@schmore That is a real bug and it is mine thank you for trying it and reporting exactly what you saw. Fixed in v1.8.3.25. You had made no error. The import gate hard-required a 4D output, so it rejected every detection model. Important for your testThe fix is in the AI service, not the plugin DLL. Pull the image again: docker pull kuscheltier/jellyfin-ai-upscaler:docker7-cpuThen: import Still worth reporting backI have no detector weights and no dog here, so the decoding is verified against synthesised ONNX models built to the documented spec, not against real ones. What I would like to know:
Six tests now cover import → load → mask as one path. I checked they reproduce your report: against the old gate five of them fail with your exact message. |
|
@schmore Both of those are real bugs, both were mine, and both are fixed in v1.8.3.28. Thank you for the detail — the second one in particular I would not have found on my own. "The status would stick on a previous model name"It was telling the truth, which is what made it confusing. The button POSTed nothing at all: the endpoint read the model id from the saved configuration. So typing a new id and pressing Load Detection Model loaded the previous model, and the status honestly reported that previous model. Worse: I had written a comment right there in the code saying "the model id is read from the SAVED config server-side, so an unsaved box would load the previous model" — and then only guarded against the field being empty. I described the trap instead of closing it. The endpoint takes a "It seemed to ignore the confidence values I put in"It did, and the reason is worth knowing if you ever hit it elsewhere. The field had So 0.6 was silently rejected. My save code skipped empty strings — which normally means "the user left it blank", but here meant "the browser disliked it". The number vanished without a word and the old one stayed. Now Why the values also revertedSame root cause as the first one, and it explains the "went away and came back" part: nothing on that card was persisted until the global Save Settings, while an action button sat right next to the fields implying otherwise. Load Detection Model now saves the card first, then loads. So pressing it stores your classes, mode, padding and confidence too. On the rest of your report
Four regression tests cover these, and I checked each one fails when its fix is reverted — the first bug existed precisely because the guard I wrote tested the wrong thing. |
Uh oh!
There was an error while loading. Please reload this page.
Share ideas for new features
All reactions