Where career-ops is going #904
Replies: 12 comments 5 replies
|
Really excited about this direction. I might be reading this as more of a recommendation system than you're envisioning — correct me if I'm off base — but if so: Cold start seems like the central tension — how are you thinking about that? The value only compounds once there's enough signal density, but early users are contributing to a mostly empty pool. Are you envisioning similarity primarily across users (people like me preferred these role types) or across jobs (this listing is structurally similar to these others)? Or both layers? And is the shared layer more of a live lookup at evaluation time, or an async enrichment that gets pulled periodically — like a community-built signal that supplements your local scores? One thing I'm really hopeful about with more signal density: role title normalization across companies. "Senior MLE at Stripe" vs "Staff Engineer, ML Platform at Meta" can be the same job at the same level, but you'd never know from the titles alone. In aggregate, community evaluations could make that mapping transparent in a way no single person's search ever could. The other angle that feels natural here is recruiter and networking signal — not just what jobs exist, but who's actually responsive, which roles are moving fast, which teams are actively hiring vs. just collecting applications. That kind of meta-signal is almost impossible to build alone but emerges quickly from a community. |
|
Love the "shared map" vision. One thing I'm struggling to understand is where the network effect comes from given how personalized career-ops is today. From what I can tell, the evaluation is heavily based on a user's CV, career story, preferences, strengths, and goals. Metrics like CV match, level fit, and career alignment seem inherently profile-dependent. Even if Person A and Person B are targeting the same role, A's evaluation is still being computed against A's profile. B would still need to run the evaluation against their own profile to know whether it's a good fit. That makes me wonder whether the real shared value comes less from the evaluations themselves and more from the metadata around them: company research, recruiter responsiveness, compensation intelligence, hidden requirements, posting legitimacy, application outcomes, etc. Also, if human evaluations become part of the signal, how do we prevent subjective opinions and biases from being amplified across the network? Curious how you think about those tradeoffs, because right now it feels like the most valuable outputs of career-ops are highly personalized, while the most shareable outputs are the job and market-level insights around them. |
US user willing to supportThis is really great. I love what you've accomplished and your vision for expanding this open source automated job search tool. I'd be happy to help from a US market and English language perspective. I'm a product leader and data scientist, so those would be the roles that I would be searching for (if/when seeding the shared role space with data) and/or pulling from shared data. Recent Usage Limit IssuesSharing data to solve "re-evaluating the exact same listings, in parallel, blind to each other" is a great idea to solve recent Claude Code usage limits I've been running up against. Right now, I actually am mostly unable to use the tool (very recently). I'm on the $20 pro plan and hesitant to upgrade to the $100-$200 Max plan. Challenges Sync'ing Repo Changes / UpdatesI haven't yet been able to keep my "personalized repo" updated with new changes you have been on-goingly making. The reason for this is that I wanted to make changes to some of the common / shared files, so I'm concerned that if I update then my changes will be overwritten. I should probably do a merge and inspect conflicts but I haven't had time to do this. Or alternatively, send my changes as feature requests for all (haven't had time for this yet). I'm also not a "software engineer" by training (my coding experience came through data science work), so merges are not my strength (lol). The "outdated common code base" is probably an issue for me as you evolve the repo in the direction you've described (probably a challenge I can over come though). Some questions to consider:
I'm asking because this level of scale is likely important to being able to consistently accumulate enough data coverage that makes the shared / available job descriptions meaningful / optimal (recency is key here). I've noticed that sometimes the websearch finds stale job roles and pulling from a shared cache of data could amplify this. Let me know if I can be of any support. Jeff |
|
Picking up on @ManishVangara's point about what's actually shareable — I think the answer is behavioral signals, not fit scores. I have two open PRs already moving in this direction: the red-flag detector from transcript signal (#1233) and the behavior tracking work. The core idea: transcript signal — interviewer latency, process disorganization, no-show reschedules, contradictory role descriptions — is relatively objective and company-attributable. It's not "this role scored 3.8 for me," it's "this recruiter cold-called without notice, rescheduled twice, and had no direct extension." That kind of signal aggregates cleanly across users regardless of their personal profile. If the shared layer is designed to receive this kind of input — anonymized, company-keyed behavioral observations rather than personalized evaluations — it starts to look a lot like a crowdsourced 996.ICU but evidence-grounded and transcript-derived rather than manually reported. (996.ICU is Anti-996 licensed and China-only so direct reuse is off the table, but the model is the right reference point.) One structural suggestion: region-scoped lists rather than one flat file. Each entry company-keyed with the behavioral signal that triggered it (sourced from transcript analysis, not subjective opinion), so the data is auditable and the bar for inclusion is consistent across regions. Flagging since the PRs are already open — happy to extend the red-flag work in this direction if it fits the roadmap. |
|
Sorry for the slow reply, @jeffallen007 — the plugin-system launch ate the last few weeks. Thanks for this: a US product-leader + data-scientist perspective is exactly the kind of seeding help a shared role space will need, and the usage-limits pain you describe (everyone re-evaluating the same listings, blind to each other) is the same one pushing this direction. Where things actually stand: the foundations are landing as open, local-first pieces first — the plugin system (opt-in integrations, BYO keys, never auto-submit) shipped in v1.15.0, and the community is converging on what is even shareable (see @Schlaflied's comment below on behavioral signals vs. fit scores — worth your read). The opt-in rails for any shared space don't exist yet; when they do, early seeding from the US side will genuinely matter and I'll flag it in this thread. If you want to get your hands dirty before that: the scanner-provider layer (#230) is where US-market coverage lives today, and provider PRs are the fastest path to shaping what data the ecosystem sees. |
|
@Schlaflied this is the most useful framing this thread has produced. A fit score is a function of you — sharing it leaks your profile. Interviewer latency, no-show reschedules, contradictory role descriptions are facts about them — observable, company-attributable, candidate-anonymous. That difference is exactly what separates shareable from not. It also composes naturally with your red-flag detector work (#1233). No architecture commitments yet — but "objective, company-attributable, candidate-anonymous" is the right test for anything that would ever cross a wire, and I'm adopting it as the working filter for this thread. |
|
Turned this into a formal RFC rather than leaving it as a buried comment: #1506 — proposes a concrete schema for the objective/company-attributable/candidate-anonymous filter @santifer adopted above, grounded in the two things already shipped/in-review in this repo (#1233's signal taxonomy, #1466/#1467's process-friction tag) rather than inventing a new one from scratch. @drkalexander1 @ManishVangara @jeffallen007 — tagged you all in the RFC with a specific open question each, since your comments above raised exactly the questions this schema needs to answer before it goes further. Would genuinely appreciate a look. |
|
Disclosure: I run freehire.me, an open-source job aggregator. Not pitching — contributing measurements, because we shipped the exact filter @santifer adopted above and it's been in production for a few days: ghost-job signal. The schema, for #1506. Four criteria, two tiers, two gates:
On cold start (@drkalexander1): across our full catalogue, plenty of postings picked up a structural stamp, only a tiny fraction converged on two criteria, and nothing reached three of four. The outcome tier has never fired — not enough distinct people per posting to clear the witness gate. Structural signal scales instantly and buys little; outcome signal is the valuable half and starts empty. Cold start isn't uniform — it lands entirely on the half that matters. Warning for the company-keyed list (@Schlaflied): in isolation, Wording: the word "ghost" never reaches our UI — a hedged label, an N/4 scale, and a checklist showing which criteria have no data. "Open 240 days, found only on an aggregator" is observable; "ghost job" is a claim about intent. Separately and much smaller: #2350 proposes freehire as a read-only, zero-auth scanner provider per #230 — no credentials, no outbound sync. |
|
Late reply, @strelov1. Thanks for the disclosure-first approach: that is exactly how we ask operators to show up here, and you did it unprompted. The production data is interesting, particularly the point that convergence rather than curation resolved the staffing confound. The schema work continues in RFC #1506 on its own schedule, and #2350 rides the normal provider-review track like any other source. |
|
Freshness/staleness for the shared pool (whenever #904's service actually gets built) — floating a design idea, not proposing to lock anything in. @santifer flagged this in #1506's round 2 as part of why the pool can't live as a git repo: "companies fix their processes — a 6-month-old friction signal is misinformation with good intentions." @drkalexander1 separately flagged the tension between coarsening I don't think a hard TTL (delete after N months) is the right shape. The design problem this whole thread keeps coming back to is cold start — @strelov1's production data showed most companies never clear even a 2-criterion bar. A hard expiry throws away exactly the scarce signal that makes low-density companies usable at all, right when the thread has spent months establishing that scarcity is the actual enemy here. Alternative: keep One thing this does NOT fix, and I don't think anything does: Curious whether decay-weighting breaks anything the trust-model/Sybil discussion depends on, or whether it's orthogonal to it. |
|
Late but real reply, @Schlaflied: this is the right way to think about the freshness tension. Deletion at the storage layer answers 'is this signal old' by destroying the answer to 'was this company ever observed at all' — and the whole thread agrees scarcity is the enemy. Weighting at the consumption layer keeps both truths: observedAt stays data (as #1506 already has it), and how much an old signal counts becomes a consumer decision instead of a storage death sentence. Noted for when the service gets built — numbers like 6/18 will be measured then, not guessed now. Keep these coming. |
|
Not quite sure where my experience fits best yet. My background is heavily infrastructure — most of the last 13 years cloud-focused. If the shared layer is going to run as a real service, there are a lot of options for doing that well without the bill getting away from you, and doing it in a way that's AI-friendly rather than fighting the tooling. Happy to dig in there if it's useful. The discussion also surfaced something I can't fully quantify yet, and it's where my other experience comes in. I've done a lot of hiring, not just applying — and I keep coming back to: how do we get hiring organizations, or the talent-acquisition firms they contract, to see this process differently? We all know what job search and talent search have become. A role collects 500–1000 resumes within hours, most of them not a fit. As a hiring manager it's a slog, and if you're lucky there are a few quality candidates at the bottom of it. As an applicant you're one of those 1000, unlikely to survive an ATS whose keyword matching was designed around the turn of the century. That asymmetry is why we're all here building templates and workarounds to make applying to a Workday or Greenhouse posting less painful — their "AI parsers" are the poster child for brain-dead applications of AI. Tempting to just go start a competitor (Workhouse? Greenday? 😄). Consider what the current shape actually asks of a candidate: create a Workday account per company, re-key every piece of information you already have in a structured file, and then watch it vanish into a void you never hear back from. And LinkedIn, which might have been the counterweight, has become a social media feed with the spam and scams to match. The thread has done good work on what's shareable candidate-side (the objective / company-attributable / candidate-anonymous filter is the right test). What I don't see anyone asking yet is whether any of this signal is worth anything to the other side of the table — the hiring manager drowning in the same slop, from the opposite direction. Both ends of this market are miserable and neither can see the other. I don't have a proposal, just a strong suspicion there's something there. |

Uh oh!
There was an error while loading. Please reload this page.
Right now, every one of us runs this search alone, in the dark. You scan boards, evaluate roles, and burn tokens doing it — and across the community, thousands of people are spending those same tokens re-evaluating the exact same listings, in parallel, blind to each other. The role you just discarded as a bad fit is someone else's dream job — and that signal just evaporates. We're all solving the same problem from zero, over and over.
That's the thing worth fixing: none of us has visibility into what the rest of us already learned.
So here's the direction. The open-source career-ops stays exactly what it is — open, local, and yours. On top of it, there's an optional, opt-in shared layer in the works: a way for the community to stop re-doing each other's work and start seeing the job market together — better matches, far less wasted effort and tokens, for everyone who chooses to join. Less "everyone alone with a flashlight," more "a shared map."
What will never change, no matter what we build:
The shared layer runs as a separate service (it takes real infrastructure to operate), so the core stays light and local — but it's built for this community, not instead of it.
Want to help? Scanner providers, new languages and markets, CLI support, docs and fixes are all welcome in the core. Bigger ideas — automation, integrations, the shared layer — drop them in the comments here and let's shape them together.
All reactions