Skip to content

NOOA 0.0.9

Latest

Choose a tag to compare

@sklinglernv sklinglernv released this 18 Aug 14:57
54e8fa2

What's Changed

  • docs(cybergym): simplify README into a script-driven walkthrough by @sklinglernv in #60
  • feat(cli): let external packages add nooa subcommands via entry points by @sklinglernv in #62
  • docs(nooa_memory): fixed broken tests/memory/ link in README by @shivaperumalsamy in #64
  • Docs/cybergym cleanups by @nv-pberard in #63
  • test(tracing): mock urlopen as a context manager in flush_pending test by @Yeriahz in #55
  • fix(nooa-bench): always persist traces and dump a local trajectory by @rdasilveiracabral in #69
  • fix(tracing): authenticate to remote viewers, and write traces where they survive by @rdasilveiracabral in #71
  • fix(examples): support commit SHAs for git_ref in the Harbor adapter 🤖🤖🤖 by @kmad in #70
  • refactor(nemo-relay)!: rename nemo_flow_* to nemo_relay_* by @sklinglernv in #61
  • fix(atif): emit real UTC step timestamps instead of local time by @sklinglernv in #59
  • fix(nemo-relay): bound dependency, migrate intercept API, fix concurrent scopes by @sklinglernv in #58
  • fix(viewer): make the frontend build hermetic and actually verify dist/ by @sklinglernv in #66
  • Honor canonical viewer plugin attribute 🤖🤖🤖 by @nirosen in #56
  • test(memory): assert owner gating blocks relay through foreign memories by @Yeriahz in #73
  • Fix 2 high-severity npm advisories in the viewer frontend by @rdasilveiracabral in #74
  • test: fill signal, sandbox, and TaskWrapper edge-case stubs by @skundu42 in #32
  • feat(strategy): allow a callable for @strategy(llm=...) by @sklinglernv in #81
  • fix(mcp): keep environment placeholders literal by @furgalep in #79
  • fix: preserve sandbox state after recoverable exits 🤖🤖🤖 by @jmolz in #42
  • ci: run sandbox containment tests, split off the sandbox marker by @alessiodevoto in #88
  • fix(ci) - ripregrep skipping by @alessiodevoto in #89
  • chore(release): codify the release process in a single script by @sklinglernv in #91
  • docs(examples): add Ollama and vLLM to the model selection snippet by @Yeriahz in #83
  • feat(tracing): add portable journal file exporter by @sklinglernv in #116
  • docs(nooa_memory): correct the relative depth on the quickstart link 🤖🤖🤖 by @Hotragn in #103
  • Remove duplicate pytest-cov dependency by @vikky781 in #133
  • Update stale nemo-oo-agents package references to nooa by @vikky781 in #134
  • feat: add shared interactive session foundation by @furgalep in #82
  • fix(skills): load and reload external skills per agent by @furgalep in #127
  • feat(cybergym): add portfolio-based NOOA agent by @sklinglernv in #145
  • fix(viewer): restore a working auth configuration for UI and ingest by @sklinglernv in #112
  • fix: Remove stray -e artifact from .gitignore by @andrewwhitecdw in #141
  • Fix memory store SQLite thread safety by @cpakkamisaac-sae in #131
  • docs(storage): correct the StorageManager snapshot example to the real API 🤖🤖🤖 by @shsagnik in #139
  • Add LLM routing metadata to generation observability by @cpakkamisaac-sae in #39
  • fix(release): use Git excludes for SPDX scan by @sklinglernv in #157

New Contributors

Full Changelog: v0.0.8...v0.0.9


🧪 Capability Test Results

⚠️ Review required — Stable tier at 93.8% (floor: 60%), 0 collapse(s), 3 new error type(s), 0 removed test(s)


92.4% 🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦🟦⬜⬜ +0.2%

1286/1392 tests passing (+3 from v0.0.8)

Metric v0.0.8 This release Change
Tests Passed 1283/1392 1286/1392 +3 ✅
Success Rate 92.2% 92.4% +0.2% ✅
Collapsed tests 0 0 ➖
New error types 3 3 ❌
Output Tokens 1,231,645 1,187,310 -44,335
Total Tokens 10,938,221 10,800,970 -137,251
📊 Per-tier breakdown
Tier v0.0.8 This release Change Expected
Frontier 119/144 (82.6%) 115/144 (79.9%) -4 / -2.8% ❌
Stable 1164/1248 (93.3%) 1171/1248 (93.8%) +7 / +0.6% ✅ ≥60%
❌ New error types
Test Model This release Baseline
refinement nemotron3-nano-30b ExecutionError 0 before
task_decomposition nemotron3-nano-30b SyntaxError 0 before
truncation_dict_codeact nemotron3-nano-30b ValidationError 0 before
📋 Per-test breakdown
Test Status
calculate_batch_001 ❌ 6/12
calculate_complex_001 ✅ 12/12
calculate_complex_002 ✅ 12/12
calculate_simple_001 ✅ 12/12
calculate_simple_002 ✅ 12/12
construction_dataframe_001 ✅ 12/12
construction_dataframe_002 ✅ 12/12
construction_widget_001 ✅ 12/12
construction_widget_002 ✅ 12/12
context_notes_001 ❌ 11/12
employee_lookup_001 ✅ 12/12
error_recovery_001 ❌ 11/12
fast_agent_help_001 ✅ 24/24
fast_agent_help_002 ❌ 23/24
fast_agent_help_003 ✅ 24/24
fast_agent_no_help_001 ✅ 24/24
fast_agent_no_help_002 ✅ 24/24
fast_agent_no_help_003 ✅ 24/24
fast_food_cancel_001 ✅ 12/12
fast_food_cancel_002 ✅ 12/12
fast_food_order_001 ❌ 11/12
fast_food_order_002 ❌ 9/12
fast_food_order_003 ✅ 12/12
fast_food_order_004 ❌ 10/12
fast_food_order_005 ❌ 11/12
fast_food_order_006 ❌ 10/12
json_extract_001 ✅ 12/12
json_qa_lookup_001 ✅ 12/12
json_qa_lookup_002 ✅ 12/12
json_qa_reasoning_001 ✅ 12/12
json_qa_reasoning_002 ✅ 12/12
large_data_count_generated_001 ✅ 12/12
large_data_extract_generated_001 ❌ 11/12
large_data_find_generated_001 ✅ 12/12
large_data_find_generated_002 ✅ 12/12
needle_in_haystack_001 ✅ 12/12
realfmt_lower_dict_codeact_001 ✅ 12/12
realfmt_lower_dict_codeact_002 ✅ 12/12
realfmt_lower_dict_codeact_003 ✅ 12/12
realfmt_lower_dict_codeact_004 ❌ 10/12
realfmt_lower_dict_codeact_005 ✅ 12/12
realfmt_lower_dict_codeact_006 ❌ 8/12
realfmt_lower_dict_codeact_007 ✅ 12/12
realfmt_lower_dict_predict_001 ❌ 11/12
realfmt_lower_dict_predict_002 ❌ 5/12
realfmt_lower_dict_predict_003 ✅ 12/12
realfmt_lower_dict_predict_004 ❌ 9/12
realfmt_lower_dict_predict_005 ✅ 12/12
realfmt_lower_dict_predict_006 ❌ 8/12
realfmt_lower_dict_predict_007 ✅ 12/12
realfmt_lower_set_codeact_001 ✅ 12/12
realfmt_lower_set_codeact_002 ✅ 12/12
realfmt_lower_set_codeact_003 ✅ 12/12
realfmt_lower_set_codeact_004 ✅ 12/12
realfmt_lower_set_codeact_005 ✅ 12/12
realfmt_lower_set_codeact_006 ✅ 12/12
realfmt_lower_set_codeact_007 ✅ 12/12
realfmt_lower_set_predict_001 ✅ 12/12
realfmt_lower_set_predict_002 ❌ 6/12
realfmt_lower_set_predict_003 ✅ 12/12
realfmt_lower_set_predict_004 ❌ 6/12
realfmt_lower_set_predict_005 ✅ 12/12
realfmt_lower_set_predict_006 ❌ 7/12
realfmt_lower_set_predict_007 ✅ 12/12
realfmt_nested_codeact_001 ✅ 12/12
realfmt_nested_codeact_002 ✅ 12/12
realfmt_nested_codeact_003 ✅ 12/12
realfmt_nested_codeact_004 ✅ 12/12
realfmt_nested_codeact_005 ✅ 12/12
realfmt_nested_codeact_006 ✅ 12/12
realfmt_nested_codeact_007 ✅ 12/12
realfmt_nested_predict_001 ✅ 12/12
realfmt_nested_predict_002 ✅ 12/12
realfmt_nested_predict_003 ✅ 12/12
realfmt_nested_predict_004 ✅ 12/12
realfmt_nested_predict_005 ❌ 9/12
realfmt_nested_predict_006 ❌ 9/12
realfmt_nested_predict_007 ❌ 11/12
realfmt_slice_keys_list_codeact_001 ✅ 12/12
realfmt_slice_keys_list_codeact_002 ❌ 11/12
realfmt_slice_keys_list_codeact_003 ✅ 12/12
realfmt_slice_keys_list_codeact_004 ✅ 12/12
realfmt_slice_keys_list_codeact_005 ✅ 12/12
realfmt_slice_keys_list_codeact_006 ✅ 12/12
realfmt_slice_keys_list_codeact_007 ✅ 12/12
realfmt_slice_keys_list_predict_001 ✅ 12/12
realfmt_slice_keys_list_predict_002 ❌ 9/12
realfmt_slice_keys_list_predict_003 ✅ 12/12
realfmt_slice_keys_list_predict_004 ❌ 10/12
realfmt_slice_keys_list_predict_005 ✅ 12/12
realfmt_slice_keys_list_predict_006 ❌ 9/12
realfmt_slice_keys_list_predict_007 ❌ 6/12
refinement_001 ❌ 5/12
repl_exploration_001 ❌ 9/12
router_analyze_001 ✅ 12/12
router_analyze_002 ✅ 12/12
router_multi_analyze_validate_001 ✅ 12/12
router_multi_transform_validate_001 ❌ 10/12
router_transform_001 ✅ 12/12
router_transform_002 ❌ 10/12
router_validate_001 ✅ 12/12
router_validate_002 ❌ 11/12
sentiment_batch_001 ❌ 6/12
sentiment_single_001 ❌ 10/12
sentiment_single_002 ❌ 10/12
sentiment_single_003 ❌ 11/12
structured_combined_extraction_001 ✅ 12/12
structured_combined_extraction_002 ✅ 12/12
structured_combined_extraction_003 ✅ 12/12
task_decomposition_001 ❌ 11/12

v0.0.8 → 54e8fa23 | 4 models × 3 runs | both arms run fresh