KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models
KGCE is a Python - centric framework designed to construct benchmark environments for Large Language Model (LLM) agents, with a specific focus on cross - platform educational agent benchmarking, integrated with multimodal language models.
🌐 Cross-Platform Educational Task Support
- Facilitates cross-platform execution (Windows, Android, educational tools) for synchronized educational workflow tasks.
⚙ ️Knowledge - Augmented Architecture
- Integrates domain-specific knowledge via a structured JSON base, handling closed-domain educational software (e.g., XiaoYa Intelligent Assistant) for accurate interactions.
📐 Dual-Graph Evaluation
- Employs a Dual-Graph Evaluator to capture educational criteria with fine-grained metrics for precise task assessment.
- Python 3.10 or newer
pip install requirements.txtAll datasets and experiment code are in kgce-benchmark directory. run
poetry run python -m kgce-benchmark.main --model qwenvl --model-base-url https://dashscope.aliyuncs.com/compatible-mode/v1 --env android --model-api-key XXX --task-id 069x