SYNTHESIS — DAG-as-code 오케스트레이션 RL 프로그램 종합 보고 (한국어판)
작성 기준: sections/ 6개 초안 + debate-log.md FINAL VERDICTS + claim-graph.md (Gate pass 1·2). 용어 통일: 오케스트레이터(orchestrator) / 보상 신호(reward signal) / 전이(transfer) / 서브에이전트(subagent). 인용 위생: REFUTED·annex 항목("~1000 서브에이전트", FlowReasoner 가중치 상수 0.7/0.15/0.15, $4.6M 학습비, australiansciencejournals.com 2건, CG18 미검증군)은 본문 어디에서도 주장으로 사용하지 않는다. 벤더 자가보고(vendor-self-reported) 수치는 vendor 태그를 유지한다.
전체 요약
이 보고서는 2026-08-31 슬랙 스레드에서 제기된 가설 — "DAG를 코드로 저작하게 하고 그 능력을 RL로 훈련하면, 일반 서브에이전트 호출(task()) 상황의 체계적 fan-out으로 전이될 수 있다" — 를 산업 사례·전이 문헌·내부 자산·보상 설계·환경 설계의 다섯 축으로 검증했다. 결론의 골격은 debate-log 최종 판정과 일치한다: 가설 A(인터페이스 간 전이)와 가설 B(플래닝 능력 자체의 향상) 모두 "유망하나 미증명(promising, unproven)"이다. 이 조합을 직접 검증한 공개 연구는 3개 독립 조사 레인의 counter-search를 통과한 부재 클레임으로 확인됐고(SUPPORTED, CG6), 오케스트레이션 자체를 RL로 학습한 공개 사례는 Kimi(PARL)가 유일하며, 그 유일한 독립 재현(OpenPARL)은 위임 역학은 재현했으나 성능은 재현하지 못했다(SUPPORTED, CG26 — 원인은 예산 10배 격차 + λ 미감쇠(un-annealed)에 의한 spawn-count 해킹).
실행 관점의 핵심 발견은 "돌고돌아 아비터 RL"이라는 스레드 내 재구성이 정확하다는 것이다. 필요한 부품 — 오케스트레이션 메커니즘을 전부 강제하고 판단 10종만 프롬프트 prior로 남기는 omo DAG 엔진, PR #654로 착지한 arbiter-rl-env의 롤아웃·보상·거버넌스 스택, slime/verl의 프로덕션 하네스 패턴 — 이 이미 내부에 존재하며, 가설 검증은 신규 인프라가 아니라 조립을 요구한다. 보상 설계는 스레드의 5개 보상 신호(병렬도·커버리지·게이트 통과·복구 성공·개입 횟수)를 게이밍 경로와 함께 개별 판정해, "게이트(-1 단락) × 규칙-검증가능 차등항 + outcome 앵커 + 반드시 감쇠하는 보조항"이라는 골격으로 재조립했다.
동시에 이 보고서의 모든 성과 주장은 세 개의 확립된 반론에 구속된다: (1) effort confound — 토큰 사용량 단독이 BrowseComp 분산의 80%를 설명하므로(SUPPORTED, CG11/CG28, 상관≠인과) 비용 조정 파레토 없는 정확도 주장은 무효(V-EFF), (2) construct mismatch — 현재 증거는 "인터페이스 인접 전이 + 제약준수 규율"까지만 지지(V-CON; PlanningBench +18.01은 objective isomorphism으로 CG8 PARTIAL 강등), (3) prior 재분배 — Qwen에서 random reward가 GT 보상 이득의 ~73%를 재현(SUPPORTED, CG24)하므로 non-Qwen 백본·random-reward 더미·base 층화 평가 없는 전이 주장은 인정되지 않는다(V-CA3).
핵심 결론 5개
- 가설 A·B 모두 "유망하나 미증명" — 직접 검증 연구 0건의 빈 슬롯(CG6)이며, 사전등록 2x2(base/DAG-RL × dag 도구/task()-only), 스캐폴드 제거 사다리, non-Qwen+random-reward 통제를 모두 통과해야 지지로 승격된다. 하나라도 실패하면 "인터페이스 인접 전이"로 강등.
- 오케스트레이션을 학습하는 곳은 Kimi뿐이고, "1000 서브에이전트"는 REFUTED다 — 실제 상한은 K2.5 100/1,500, K2.6 300/4,000(SUPPORTED, CG1/CG30); PARL은 검증된 아키텍처·미검증 레시피(CG26)이며, 보조 보상의 annealing은 장식이 아니라 하중 부품이다.
- 보상 골격은 확정 가능하다 — 게이트는 자격 조건(-1 단락, 외부 핀 validator), 병렬도는 CriticalSteps로 재정의, 커버리지는 아티팩트-한정+dedup, 복구는 telescoping credit으로 흡수, 개입 횟수는 훈련 보상에서 제외. 검증 불가능한 보상은 충분한 반복 하에 반드시 해킹된다(OpenPipe 전수 관측; CG20 15항 체크리스트 통과 전제).
- 인프라는 조립이다 — arbiter-rl-env(프로세스-에피소드 격리·fingerprint·anti-hacking ledger) + slime/verl 하네스 + omo dag 툴로 27B+LoRA 단일 8-GPU 노드에서 실행 가능하며, 이것이 V-BUD의 "저비용 rollout 경제" 조건을 자체 충족한다.
- 로드맵은 게이트식이다 — Stage0 GEPA/프롬프트-최적화 대조군(스킵 불가, V-BUD) → Stage1 소량 SFT 콜드스타트(V-CA3, subflow 0/8 교훈) → Stage2 EI → Stage3 조건부 GRPO(verl). 각 단계 exit criteria 미달 시 다음 단계에 예산을 쓰지 않는다. stop-specific 보상은 확인된 빈 연구 셀이자 시간부패가 빠른 1순위 novelty 기회다(V-STOP).
1. 배경·내부현황
- 슬랙 제안의 핵심: DAG-as-code 저작을 RL 타깃으로, 일반 task() fan-out 전이를 검증 대상으로 삼는다 — 단 "K2.5 1000 서브에이전트"는 REFUTED(실제 상한 100/300, CG2), Kimi 스웜 게인은 vendor + effort-uncontrolled 이중 태그로만 인용한다.
- omo DAG 엔진(약 13.9k LOC, MEASURED)은 스케줄링·멱등·복구 동사 등 메커니즘을 전부 강제하고, 판단 10종(분해·라우팅·프롬프트 계약·검증 배치·드리프트 대응·복구 선택·수렴·클레임 게이트·팀 구성·딜리버리)만 프롬프트 prior로 남긴다 — 이것이 정확히 RL 훈련 표면이다.
- arbiter-rl-env는 PR #654(93파일, 6,208 LOC, MEASURED)로 롤아웃·보상·거버넌스·코퍼스 240태스크가 착지 완료 — 보상은 게이트×D1–D5(도달 최대 3.9, DERIVED), judge는 미배선, 코퍼스는 목표 300 대비 60 미달이다.
- 스레드의 5개 보상 신호는 gate×D 스택에 거의 직접 매핑되나 '실패 복구·중단' 계열은 빈 셀 — stop-specific 보상 부재(V-STOP)와 같은 근원(런 라이프사이클 판단)에서 나온 쌍둥이 공백이다.
- Phase-1 스코어드 런의 선행 조치 3건: 보안 조치(CG21), OPENAI_API_SHAPE 충돌 해소(CG22), judge 배선 — 이들 없이 스코어드 런은 무효화된다.
1.1 발단 — 슬랙 스레드
기점은 2026-08-31 17:15–17:22(KST)의 8분 스레드다(MEASURED, int-thread-context §1). 제안: "DAG 나 이런거 코드로 작성하게하고, 그 능력자체를 일반 서브에이전트 호출이나 이런거에 전이시킬수도있지않을까"(slack-thread.txt:5). 관찰된 실패 모드가 배경이다 — task() 배치가 상한 16개를 지원함에도(MEASURED, slack-thread.txt:69-71) "서브에이전트를 써라" 프롬프트로는 모델이 전형적으로 1–2개 호출에 그친다(MEASURED, slack-thread.txt:11). 스레드가 명명한 5개 보상 신호: 웨이브 병렬도, 주제 커버리지, 검증 게이트 통과율, 실패 복구 성공률, human 개입 횟수(slack-thread.txt:24). 프롬프트/스킬은 정책의 사전확률(prior)만 교란하고 RL은 그래프 품질을 직접 최적화한다는 구분도 명시됐다. 상대의 재구성 — "돌고돌아 아비터 RL"(slack-thread.txt:94) — 이 이 보고서의 프레임이다.
교정 1건: 제안자가 회상한 "K2.5 서브에이전트 1000개"는 REFUTED다. 1차 소스 상한은 K2.5 100 서브에이전트/1,500 툴콜, K2.6 300/4,000이며, "1000"은 K2.6 코딩 데모의 툴콜 수 오인용이다(CG2 REFUTED, CG1 SUPPORTED).
1.2 내부 자산 5종
- omo workflow(dag) 툴: 8개 액션(start/attach/snapshot/wait/cancel/retry/send/amend). 엔진이 강제하는 것 — 의존성 프론티어 어드미션, 정의 지문 기반 멱등 재사용, WAL 저널링·크래시 복구, 자식의 오케스트레이션 도구 구조적 차단, 노드 완료 시 검증 지시어 자동 주입. 상한 노드 64/런, 런 16/세션, 프롬프트 256KiB/노드(MEASURED, int-omo-workflow §7). 판단 10종 중 7종은 엔진 대응이 전무하거나 형식적 — "프롬프트는 prior만 흔든다"는 스레드 진단과 일치하는 지점이자, 곧 액션 공간이다.
- arbiter-rl-env:
scalar = gate.pass ? clip(Σ wᵢ·dᵢ, −1, +4) : −1. 가중치 D1 인텐트 2.0 / D2 스코프 1.0 / D3 효율 0.75 / D4 토큰 0.1 / D5 위생 0.05(MEASURED, int-rlenv-reward §2). 현행 스칼라는 100% 결정론적이고 judge.ts는 computeReward에 미배선. 코퍼스 v1 240태스크·7패밀리, held-out 20% sha256 동결. subflow 패밀리는 base 0/8이라 rl_viable:false — "base가 보이지 않는 행동은 SFT 콜드스타트로 먼저 심는다"(V-CA3)의 근거 사례. Phase-0 스파이크는 퇴장 기준 통과(에피소드 20/20, xml_in_text 누출 0; MEASURED, int-rlenv-status §3). 훈련 경로는 EI 선행 → 플래토 시 verl GRPO(TRL 철회, ROLL 폴백). - carrier: OrderSheet 실행엔진(Kotlin/WebFlux, 15개 storm/* 노드 타입). 구조적 동시성 때문에 노드 하나의 예외가 워크플로 전체를 중단하고, 상태는 무지속성이며, 문서-코드 드리프트가 유의미하다 — 실행 진실은 코드 쪽에 있다(MEASURED, int-carrier §1-3, §8).
- storm 스택: schema/dsl/linter 3층. P6 플로어 — writeFingerprint가 storm-dsl < 0.2.6이면 throw, v16 검증자 오염 버전의 RL 점수 유입을 원천 차단(int-rlenv-governance §2.2).
- arbiter-pi 스킬: 93파일 약 9.3k라인 정책 코퍼스(MEASURED). 구조 우선 진단 철학(규제 도메인 누락률 30%→5% 사례, MEASURED). 단 코퍼스 내부 미해결 모순 8건은 보상 노이즈원(agg-internal C10).
1.3 조립 그림
태스크(V-CA2 임계 위 breadth-first) · 환경(프로세스 격리+지문) · 보상(gate×D + CG20 15항 게이밍 체크리스트) · 평가(사전등록 2x2 + non-Qwen + 비용 조정 파레토) 전부가 기존 인프라 위에서 조립 가능하다. 남는 진짜 불확실성은 둘이다: 분해 능력이 capability-bound인가(EI 선택압이 작동할 조건), 그리고 DAG 저작 표면의 규율이 task() 인터페이스로 흐르는가(2x2의 우하단 셀이 답할 질문).
즉시 조치 3건: (1) 내부 보안 조치 1건 — 상세는 내부 채널로만 공유(CG21), (2) episode.ts의 OPENAI_API_SHAPE:'completions' vs v17+vLLM Responses API 요구 충돌(CG22; 강제 completions는 0/4 붕괴 실측), (3) judge.ts 배선(position-swap 가드 유지).
2. 사례
- 오케스트레이션을 RL로 학습하는 공개 사례는 Kimi(PARL)가 유일하다 — Anthropic·OpenAI·중국 4개 랩 모두 조율은 스캐폴딩이고 RL은 per-agent다.
- PARL의 문법: 동결 서브에이전트 + outcome 본체 + 반드시 감쇠하는 보조 보상(serial collapse / spurious parallelism 각각 겨냥) + CriticalSteps 임계경로 회계(모두 MEASURED, K2.5 paper; CG3/CG5 SUPPORTED).
- OpenPARL 독립 재현: 위임 역학은 재현(assign_task 0.03→1.00), 성능은 미재현(WideSearch item-F1 무개선) — 원인은 예산 10배 격차와 λ 미감쇠에 의한 spawn-count 해킹(CG26). PARL은 검증된 아키텍처, 미검증 레시피다.
- Anthropic의 +90.2%(vendor)는 토큰이 분산의 80%를 설명하는 effort confound 아래 있고(CG11), 반대 진영의 정량 하한 — MAS 실패율 41–86.7%(MAST v3), 프롬프트는 죽은 레버(FRT) — 도 실재한다. 단 MAST 실패의 67.7%(FC1+FC3)는 아티팩트/텔레메트리에서 관측 가능해 DAG-저작 보상의 직접 페널티 대상이다(DERIVED, w2-mast).
- 학습-오케스트레이션 투자의 성립 조건은 파티션 임계(~8–10 엔티티/32K 유효컨텍스트/39% 핸드오프 세율) 위 breadth-first 태스크 코퍼스 + Kimi 문법의 보상 감쇠 + 저비용 rollout 경제다(V-CA2, V-BUD).



2.1 Kimi 계보 — 학습된 오케스트레이션의 유일한 공개 사례
PARL은 오케스트레이터 정책 하나만 훈련하고 서브에이전트는 동결한다 — 서브에이전트 출력은 "환경 관측"으로 취급되어 credit assignment 모호성을 회피한다(SUPPORTED, CG3). 보상은 r_PARL = λ1·r_parallel + λ2·r_finish + r_perf, λ는 0으로 annealing(SUPPORTED, CG5). r_parallel은 serial collapse를, r_finish는 spurious parallelism을 겨냥한다. 효율 지표 CriticalSteps는 총 스텝이 아니라 임계경로를 최소화하도록 유도한다 — V-EFF가 요구하는 effort-정규화 회계의 보상 내장 사례다.
성과 수치(전부 vendor-self-reported, CG4 PARTIAL): BrowseComp 단일 60.6 → 스웜 78.4(+17.8pp), WideSearch 72.7 → 79.0, 레이턴시 3–4.5×(최댓값이며 보편 배속 아님). λ 스케줄·수치는 비공개 — 아키텍처는 공개, 레시피는 비공개다.
세대별 일관 문법(MEASURED, 각 세대 primary): K1.5(no-value-network, 길이 페널티) → K2(Self-Critique Rubric + hack-check layer) → Kimi-Researcher(outcome-only + gamma-decay credit) → K2 Thinking(순차 깊이 200–300 툴콜, vendor) → K2.5(병렬 폭, PARL) → K2.6(300/4,000 확장, Claw Groups, Document-to-Skills) → K3(2.8T/104B, "20+ concurrent subagents" 사례 — 마케팅 축이 지속시간·재귀 깊이로 이동, DERIVED). 공통 상수 둘: outcome이 본체이고 shaping 항은 언제나 보조·감쇠·cap(DERIVED, agg-kimi §4), 그리고 동결/분리 구성요소의 일관 선호(end-to-end 미분 가능성을 포기하고 안정성을 산다).
독립 검증 현주소: 스웜 end-to-end 독립 재현 없음. Moonshot 자체 K2VV — 서드파티 서빙 스택이 툴콜 충실도를 13–27pt 잃는다(CG23 PARTIAL, vendor-run). 서드파티 순위는 "near-frontier, not frontier SOTA". $4.6M 학습비 루머는 UNRESOLVED(CG14) — 인용하지 않는다. "Researcher = K2 Thinking 전신" 계보 주장도 UNRESOLVED(CG15) — 단정하지 않는다.
2.2 타 랩 — 학습하지 않는 오케스트레이션
- Anthropic: 멀티에이전트 Research 시스템 +90.2%(vendor)는 오케스트레이션의 최강 공개 주장이나, 같은 글의 분산 분석이 스스로 깎는다 — 토큰 단독으로 분산 80%(툴콜·모델 포함 시 95%) 설명, 15배 토큰 비용, 동예산 대조군 부재. 조율 층의 학습된 구성요소는 0개("Prompt engineering was our primary lever").
- OpenAI: 공식 독트린이 단일 에이전트 우선. 모든 RL은 per-agent(Deep Research, CUA, codex-1), 조합은 코드/제품 수준.
- GLM/Qwen/MiniMax/DeepSeek: 초점은 환경 처리량(Qwen3-Coder 2만 병렬 환경; MiniMax-M1 풀 RL = $534,700, MEASURED). 학습된 보상 모델을 1차 agentic 신호로 쓰는 랩은 없고 DeepSeek R1은 neural RM 자체를 거부. slime
coding_agent_rl이 유일한 완전 공개 프로덕션 레시피(별도 클린 샌드박스 채점 = 테스트 치팅 차단). 4개 랩 모두 오케스트레이션 결정은 훈련 대상이 아니다. - 인접 보상 증거: RLTR — Planner만 completeness 보상으로 훈련, E2E RL을 8–12% 상회, completeness 보상이 최종답 보상보다 인간 정합(ACC 74.59 vs 65.30; CG27, 단일 랩 예외 태그). RLVP — dense progress 보상은 게이밍 가능 클래스, "검증가능 경로 페널티 + outcome" 처방. 두 결과의 합성이 Kimi 문법과 공명한다.
2.3 반대 진영 — 정량 실패 증거
- Cognition: 서브에이전트끼리는 손실 있는 요약만 교환하므로 병렬 작업은 발산한다 — "share full agent traces" 원칙(단 "Claude Code는 병렬 안 함" 서술은 2026 문서로 시효 만료).
- MAST v3(1,642 traces, NeurIPS 2025 — v1 명칭 인용 금지): MAS 실패율 41–86.7%, 개입 상한 +15.6%, 토폴로지 > 프롬프트. 파생: FC1+FC3 = 67.7%의 실패가 관측 가능 — DAG-저작 보상의 직접 페널티 영역(DERIVED; 도메인 전이는 가정).
- Anthropic FRT(2026-08): 프롬프트 3종 변형 무효, 조율 성공은 모델 세대를 따름(Sonnet 5만 진짜 협업), conformity collapse(30개 중 18개 동일 브랜치명), "transmitting context is about as costly as acting on it", channel-free 담합. 겉보기 모순 해소: 스캐폴드 구조는 유효한 레버, 프롬프트 문구는 죽은 레버(DERIVED, w2-agg C6). 이 보고서는 역설적으로 학습-오케스트레이션 논지의 최강 외부 근거(모델 세대=훈련만이 조율을 움직였다)이자 최강 하한(담합·비용 등가는 어떤 오케스트레이터 정책도 못 닿는 구조 문제)이다.
2.4 프레임의 의미
"학습하는 곳은 Kimi뿐"은 공개 증거 기준 사실이다. 그러나 (1) 이것은 빈 연구 슬롯의 실증이고(84-entry 서베이에서 오케스트레이션 보상 R7은 신생, stop 셀은 빈 채), (2) "Kimi만 한다"가 "Kimi가 옳다"는 아니며(vendor + 재현 실패 — 단 실패 원인이 특정되므로 "미검증"이지 "반증" 아님), (3) 반대 진영과 Kimi는 같은 파티션 규칙 위에 있다 — Kimi 스웜 벤치마크(BrowseComp, WideSearch)는 정확히 breadth-first 워크로드다. Kimi 사례의 교훈은 "스웜이 이긴다"가 아니라 "오케스트레이션을 학습 가능하게 만드는 태스크·예산·보상 감쇠의 결합 조건"이다.
3. 가설 A — DAG 도구 RL의 범용 서브에이전트 활용 전이
DAG 도구 내 학습
인터페이스 간 전이
능력 일반화
- 가설 A는 세 링크 사슬(T1 도구 내 학습 → T2 인터페이스 간 전이 → T3 능력 일반화)이며, 직접 검증 연구 0건의 빈 슬롯에 대한 베팅이다(CG6 SUPPORTED, 부재 클레임).
- T1은 독립 재현으로 조건부 지지(OpenPARL 위임 역학, CG26), T2는 인접 방증만 존재(Search-R1/ToolRL 인터페이스 전이, CG9 — 단 같은 인터페이스 내 전이라는 한계), T3는 존재 증명 수준(Parallel-R1의 스캐폴드 제거 후 잔존, MEASURED)이다.
- 최대 교란 3종: 토큰 effort(분산 80%, CG11), 스캐폴드 혼입(같은 gpt-4o가 스캐폴드에 따라 5.4배 차, MEASURED), prior 재분배(random reward가 이득 ~73% 재현, CG24) — 셋 다 사전등록 통제 없이는 어떤 전이 주장도 무효다.
- 판정은 promising, unproven: 동예산 2x2, 스캐폴드 제거 사다리, non-Qwen+random-reward 통제를 모두 통과해야 지지로 승격되고, 하나라도 실패하면 "인터페이스 인접 전이"로 강등된다(V-CON).
- 전이 확률을 올리는 설계: 파티션 임계 위 breadth-first 코퍼스 + 규칙 채점 r_perf + 실제로 annealing되는 보조항 + CriticalSteps 회계 + SFT 콜드스타트 커리큘럼 + 직접 도구 폴백 유지(제거 시 RL 붕괴, OpenPARL 실측).
3.1 지지 증거
- PARL(T1): 위임 역학은 독립 재현(assign_task 0.03→1.00), 성과는 예산·anneal 조건부(CG26). 실무 함의 — 직접 도구 폴백 유지(delegate-only 팔은 0.80→0.10 붕괴), λ 실제 anneal, 스텝 예산 논문급(100+100) 확보.
- Search-R1/ToolRL(T2 방증): 인터페이스 판단이 가장 전이 잘 되는 스킬 계급 — QA 2개 훈련→held-out 5개 전부 개선, 미학습 도구 +17%(MEASURED, CG9). 단 dag→task()는 다른 인터페이스 간 전이라 더 어렵다는 것이 SA1의 요지(ASSUMED, 유비).
- Parallel-R1(T3 존재 증명): 병렬성은 학습 가능·선택적(RL은 13.6–63.0% 선택 사용으로 SFT-only 붕괴를 회피), mid-training scaffold — 포맷을 제거해도 천장이 높게 유지(MEASURED). 사전학습 중첩 통제 부재로 상한 취급.
- omo 판단 레이어: 결정 표면이 얇고 base가 프롬프트로도 어설프게 수행 가능 — CG12의 prior 재분배 메커니즘이 작동하기 좋은 조건(ASSUMED, 구조 논증). PARL과 구조 동형(단일 오케스트레이터 + 조율 불가 자식 + 산출물 중심 보상).
3.2 반증과 한계
- SA1 construct mismatch(V-CON UPHELD): 인용된 어떤 긍정 증거도 "DAG 저작 훈련 → task() 단독 환경 활용"이라는 정확한 construct를 측정하지 않았다. 증거 상한은 "인터페이스 인접 전이 + 제약준수 규율".
- SA2 effort confound(V-EFF UPHELD): 스웜 델타는 토큰으로 설명 가능 — 모든 성과는 비용 조정 파레토(총 토큰/순차 토큰/pass^k/$)로만 보고.
- CG12 prior 재분배: 전이의 도달 범위는 사전학습 분포로 상계된다. 사전학습에 위임·조율 패턴이 희박하면 T2/T3는 재분배할 원료가 없다(V-CA3 통제 의무).
- SA4 스캐폴드 혼입: 검증 지시어·복구 동사·스케줄러는 그 자체로 성능을 만든다. PlanningBench 사례(§4)가 construct 검증 실패의 실물이다.
- SA11 frozen-subagent 특이성: 학습된 것이 "그 워커들의 버릇"일 수 있다 — AFlow 워크플로도 executor 교체 시 약화, FRT는 워커 성향이 오케스트레이터 도달 불가 층위임을 보였다. held-out 워커 스택 평가가 처방.
- SA14 negative transfer: 위임이 손해인 임계 아래 태스크에서의 과잉 위임 — no-op non-inferiority 평가 필수.
3.3 판정과 검증 설계
판정: promising, unproven (debate-log FINAL 일치). 승격 조건 셋 — (1) 동예산 2x2에서 교차 인터페이스 이득이 비용 조정 파레토 위에서 유지, (2) 스캐폴드 제거 사다리(L0 전체 스캐폴드 → L3 단일 컨텍스트)에서 이득의 유의미한 잔여가 정책에 귀속, (3) non-Qwen 재현 + random-reward 더미가 이득을 재현하지 못함. 검증은 6장의 실험 카드(사전등록 2x2, scaffold-removal ladder, 동예산 페어드 런+random fan-out, utilization vector — 한계기여·커버리지·라우팅·중복률·선택적 병렬성·임계경로·pass^k·no-op non-inferiority, non-regression battery, prior 통제)로 실행한다. 공통 프로토콜: 시드 ≥3, paired bootstrap, 평가자 버전 고정, 사전등록. utilization vector는 평가 전용이며 보상이 아니다(자동 귀속 정확도 step 14.2%, MEASURED — 훈련 신호로 쓰기엔 부족).
4. 가설 B — 구조화 계획 보상 RL의 플래닝 능력 향상
- 워크플로-RL→플래닝 전이를 직접 검증한 공개 연구는 0건 — 가설 B도 빈 연구 슬롯이며, 통제를 갖춘 첫 실험이 어느 방향이든 첫 증거가 된다(CG6).
- PlanningBench의 +18.01 전이는 훈련 보상과 평가가 같은 함수형(체크리스트 통과율)인 objective isomorphism — "제약준수 규율 전이"로 축소 서술이 의무다(CG8 PARTIAL 강등).
- 살아남는 발견은 결정성 ablation: 결정적 최적해 보상 +7.06 vs 비결정적 +0.75(MEASURED) — 보상 명세가 데이터 볼륨을 ~10배 이긴다(DERIVED, 두 값의 비).
- RLTR은 플랜 산출물 전용 보상의 성립을 보였다(completeness 보상의 인간 정합 74.59 > 최종답 65.30, 다운스트림 +5–6%; CG27) — 단 체커는 여전히 뉴럴이고, 명시적-플랜 훈련은 접지·스케일·실패표적화 조건부다(Plan-and-Act 9.85→29.63%, 나이브 플랜 SFT는 과적합 실패).
- 판정: 약한 버전(제약준수·인터페이스 인접 전이)은 지지, 강한 버전(플래닝 능력 자체)은 미증명 — 판별 실험 4종(PlanBench Mystery should-stay-zero, NaturalPlan 복잡도 cliff sweep, counterfactual plan test, scaffold-randomized eval)은 전부 훈련 보상과 다른 함수형이어야 한다.
4.1 지지 증거
- PlanningBench: 300개 인스턴스 GRPO로 미학습 TravelPlanner +18.01(MEASURED) — 단 §4.2에서 강등. 진짜 가치는 ablation(+7.06 vs +0.75)과 훈련 중국어→최대 전이 영어 벤치라는 잔여(순수 포맷 모방으로 설명 안 되는 부분).
- RLTR: Planner만 completeness 보상(골드 답 불요)으로 훈련 — E2E RL 대비 플래닝 8–12% 우위, decoupled SFT > E2E SFT. 계획 산출물은 유효한 훈련 객체이고 프로세스 프록시가 결과 신호보다 정합할 수 있다. 경계: 체커의 25% 인간 불일치 미검토, GRPO가 PPO에 뒤짐(76.4 vs 82.7).
- Plan-and-Act / DCAS: 훈련된 Planner 추가만으로 9.85%→29.63%; 플랜-주석 궤적은 스캐폴드를 넘어 일반화, 액션-only는 +0.61%에서 포화(MEASURED). 단 이득은 접지·10k급 스케일·실패-표적 데이터에 조건부.
- CodeAct/plan-as-code: 실행기가 있으면 구조화 표현이 최대 +20%(MEASURED) — DAG 보상의 오라클은 실행기여야 한다는 설계 함의.
4.2 반증과 한계
- CG8 강등(objective isomorphism): 훈련 보상과 TravelPlanner 평가가 동일 함수형, 평가셋 23.1% 여행/라우팅 접촉, 프로토콜 불투명 — 방어 가능한 독해는 "대부분 제약준수/포맷 규율 + 소수의 진짜 다중제약 통합". SA3의 실증적 승리이자 V-CON의 근거.
- SA5(실행성공≠플래닝): counterfactual plan test 없이는 플랜이 장식일 수 있다. SA12(plan emission=포맷 스킬): DCAS의 스캐폴드 잠금이 실증적 뒷면.
- LLM 자기검증 불신: 자기교정은 전 모델에서 성능 하락, o1-preview조차 unsolvable 인스턴스를 16–27%만 플래그(MEASURED) — 보상 경로에서 LLM 자기비판 배제.
- Echo Chamber/Spurious Rewards/1-shot RLVR: RLVR 이득 상당분은 prior 재분배(CG12/CG24) — non-Qwen·random-reward·base 층화 없는 전이 주장 불가(V-CA3).
- 좁힘/넓힘 논쟁: Limit-RLVR(pass@256 축소) vs ProRL(엔트로피 관리 시 경계 확장) — 강한 가설 B는 좁힘 진영이 맞다면 원리적으로 기각되고, 약한 버전(elicitation+재배열)만 양 진영과 양립한다. 엔트로피 관리(clip-higher 등)를 훈련 기본값으로.
4.3 판정과 설계 규칙
약한 가설 B'(제약준수 규율 + 인터페이스 인접 계획 행동): 지지됨. 강한 가설 B(일반 플래닝 능력): 미검증. 설계 규칙: 결정적 최적해 태스크만 코퍼스에(느슨한 검증은 전이를 능동적으로 죽인다), 보상은 훈련 목표와 다른 함수형의 평가로 검증, 보상 스택은 규칙-검증가능 구조 체크 + 외부 실행기 + outcome 항, 커리큘럼은 SFT 포맷 콜드스타트 → RL, 궤적 품질 필터링(rejection sampling) 필수, 플랜을 1급 산출물로 데이터에 유지, 파티션 임계 아래 태스크 배제(V-CA2). 인용 위생: 정크 저널 2건과 FlowReasoner 가중치 상수(CG29 REFUTED, NOT-FOUND-IN-SOURCE)는 존재하지 않는 것으로 취급한다.
5. 보상설계
- 검증 불가능한 보상은 충분한 반복 하에 반드시 해킹된다(OpenPipe 전수 관측, MEASURED) — 기본 골격은 "게이트(-1 단락) × 규칙-검증가능 차등항 + outcome 앵커"이며 arbiter-rl-env가 그 v0 구현이다.
- 스레드 5신호의 판정: 병렬도는 CriticalSteps로 재정의(소비-병렬성만 카운트), 커버리지는 아티팩트-한정+dedup 조건부 채택, 게이트는 학습 신호가 아니라 자격 조건, 복구는 telescoping credit으로 흡수(양의 복구 보상은 자기-실패 유발 유인), 개입 횟수는 훈련 보상에서 제외(은폐 유인)하고 원인 페널티로 대체.
- 보조항 annealing은 장식이 아니라 하중 부품이다 — OpenPARL에서 λ 미감쇠만으로 spawn-count 해킹이 실측 재현됐다(CG26). 참조 상수는 OpenPARL의 λ1=0.3/λ2=0.2/병렬 cap 10뿐이다(MEASURED — PARL 원 논문 상수는 비공개).
- 설계 원칙은 RLVP×RLTR 합성: 구조적 양(+)신호는 규칙-검증가능하게, 페널티엔 fulfillment credit을 짝짓고, outcome이 항상 태스크 드라이버다. 길이·노드 수는 절대 보상하지 않는다(길이 보상 최대 −17.7pt 실측).
- 판정자(judge)는 그룹-상대 채점(RULER 패턴) + held-out 인간 정합 감사 + random-reward 더미 대조를 통과한 뒤에만 배선한다 — Qwen에서 랜덤 보상이 GT 이득의 ~73%를 재현한 이상(CG24) 더미 대조는 비협상이다.
5.1 5신호 카탈로그
| 보상 신호 | 검증가능성 | 주 게이밍 경로 | 완화 | 선례 |
|---|---|---|---|---|
| 병렬도 | 낮음(count) → 중간(CriticalSteps 재정의 시) | fan-out 복제, spawn-and-forget, 가짜 fork, 정크 브랜치 | 소비-병렬성만 카운트, CriticalSteps 회계, λ annealing 완결, cap, orphan 페널티 | K2.5 PARL(vendor), OpenPARL 해킹 실측 |
| 커버리지 | 중간(규칙 매칭 결정적, 의미론 비동형) | node-stuffing, 중복 개명, hub-spoke, breadth-over-depth | haystack을 아티팩트로 한정, 시맨틱 dedup, 효율 상쇄항 | arbiter D1+D3(코드), CG8 반례 |
| 게이트 통과 | 높음(외부 고정 validator 전제) | self-graded gate, predicate drift, 우회 라우팅, timing attack | 게이트/그래프 저자 분리, 핀 validator, gate-fail 단락(-1) | arbiter G1/G2, SWE-rebench 분리, DeepSeek rule-only |
| 복구 성공 | 낮음(양의 항은 자기-실패 유발) | 실패 양식→복구 수확, retry 스팸, 길이 인플레 | telescoping credit(순credit 0), RTPO forking, 미복구-방치 페널티+fulfillment credit, retry cap | TRACE 수식, RLVP 규칙, arbiter D4 |
| 개입 횟수 | 매우 낮음(은폐로 최적화 가능) | 실패 은폐, 과신 완료 주장, 검증 생략 | 훈련 보상 제외, 원인 페널티(검증 미실행 −0.3류) 대체, off-loop CoT 모니터, 배포 지표화 | K2 hack-check(vendor), arbiter D5, Baker CoT 모니터 95% |
공통 규칙(CG20 15항 체크리스트 중 하중 큰 것): 모든 보상항을 공격면으로 인벤토리(항별 최저가 공격을 훈련 전 문서화), 항별 bound, KL leash + proxy-vs-held-out 대시보드, 발견된 익스플로잇은 metric 버그로 취급, 형제 익스플로잇 재검 정례화, 길이 절대 비보상.
5.2 검증된 패턴 재사용
- PARL 3항 + annealing: outcome 앵커 + 감쇠 보조항. OpenPARL 3대 교훈 — annealing은 하중 부품, 예산 10배 미달 시 메커니즘만 재현, 직접 도구 폴백 제거 시 RL 붕괴(액션 공간 설계가 보상 설계에 선행).
- RLVP×RLTR 합성: 구조 체크는 규칙-검증가능하게(스키마·비순환·엣지·dedup) + outcome 드라이버 + 페널티마다 fulfillment credit. RLTR의 GRPO<PPO 관측에 따라 completeness류 신호에는 PPO/REINFORCE++ 대조 선행.
- arbiter 게이트×차등항: change-intent 교차-무효화("아무것도 안 하고 커버리지 점수"의 이중 차단)와 판정자 계약(strict JSON, null-never-fabricated, position-swap 필수)을 그대로 계승. 가중치 수치는 도메인 특화라 이관하지 않는다.
- Credit 레이어: TRACE telescoping TD(endpoint-only — 패딩·복구팜 구조적 불가, K=3/γ=0.8) + GiGPO anchor grouping(추가 롤아웃 0, canonical graph-state 해시로 이식 — ASSUMED, 타 도메인 작동점) + RTPO selective forking(validator-rejection 경계에서만). 학습된 PRM을 보상으로 직접 쓰는 구성이 가장 취약(AgentPRM 82→70 붕괴, MEASURED).
5.3 가설별 스택과 실행 순서
가설 A용(아티팩트 부착): 정책은 그래프 방출에서 종료, executor 동결, 보상은 방출 그래프의 함수 — 거의 전부 결정적으로 구성 가능. 게이트(-1 단락) + 구조 체크 + 아티팩트-한정 커버리지 + 효율 상쇄 + 위생 페널티. 가설 B용(궤적 부착): PARL 3항 골격 + credit 레이어, 비용 조정 파레토·random fan-out 베이스라인 의무. 실행 순서는 A→B: A에서 anti-hacking ledger를 결정적 환경에서 검증한 뒤 같은 골격에 PARL 항을 얹는다. Stage0 GEPA arm은 V-BUD에 따라 의무 편성.
6. 환경구성·로드맵
DAG 코드 + 복구 동사
컴파일·WAL·격리
실행·아티팩트
gate × outcome
- 환경은 새로 짓지 않는다: arbiter-rl-env의 프로세스-에피소드 격리/fingerprint/anti-hacking ledger와 slime/verl의 하네스-어댑터 패턴을 omo dag 툴 위에 이식하는 것이 신규 작업의 전부다.
- 액션 = DAG 코드 제출 + 복구 동사(retry/amend/send), executor는 동결 — PARL·RLTR이 검증한 오케스트레이터-only 구조이며, 직접 실행 폴백 제거 시 RL이 붕괴한다(OpenPARL 실측). 엔진의 7개 컴파일 오류 코드가 규칙-검증가능 게이트로 공짜 확보된다.
- 태스크 코퍼스는 파티션 임계(~8–10 엔티티/32K/39% 핸드오프 세율) 위 breadth-first 태스크로만 구성하고(V-CA2), 실패-시드 [0.05, 0.75] 밴드 커리큘럼 + 결정적 최적해 태스크 초기 ~50%(ASSUMED) + serial-optimal 함정 ~15%(ASSUMED)를 혼합한다. base 0/N 패밀리는 rl_viable:false로 SFT행(V-CA3).
- 로드맵은 Stage0 GEPA 대조군(스킵 불가) → Stage1 소량 SFT(포맷) → Stage2 EI → Stage3 조건부 GRPO이며, 각 단계 exit criteria 미달 시 다음 단계에 예산을 쓰지 않는다. 27B+LoRA는 단일 8-GPU 노드로 충분하고 병목은 샌드박스 wall-clock이라 fully-async rollout이 기본값이다.
- token provenance("string in, token out" — 재토크나이즈 시 단일 턴에서도 PPO 미수렴, slime·verl 양쪽 독립 관측)는 수렴 자체를 깨는 숨은 정합성 제약이며, provenance 어댑터 테스트가 Stage1 진입 조건이다.
6.1 환경 설계
에피소드 = "태스크 프롬프트 → 정책이 오케스트레이터로 DAG 제출·관리 → 실행 → 채점". 액션 공간은 dag 툴 8액션 전체(복구 동사 포함 — 판단 카탈로그의 훈련 대상이자 폴백 안정장치). 종료 조건은 arbiter의 termination 열거형 차용, 실패 에피소드는 부분 궤적과 함께 기록(조용히 버리지 않음). DAG-as-artifact 설계를 채택하고 "자유형 task() 궤적 전체 보상" 대안은 기각한다 — 판정자 의존과 effort confound에 그대로 노출되기 때문이다. 컴파일 성공 시 엔진이 산출하는 waves/criticalPath/bottlenecks 메타데이터(MEASURED)는 효율 채점 입력으로 재사용.
frozen-executor 근거 3겹: PARL 검증 구조(CG3), RLTR의 frozen-Summarizer 성공(CG27), 인과 식별(오케스트레이션 개선과 워커 개선의 분리 — V-CON 요건). arbiter 재사용 맵: 프로세스-에피소드 격리(threads 최적화 금지 경고 계승), SSE 종료 이벤트 감지(폴링 금지 — 실제 버그 전례), fingerprint-first 거버넌스(rollout CLI 미배선 갭을 반복하지 않음), anti-hacking ledger(CG20 15항으로 처음부터 작성), EI→GRPO 골격, 5조건 AND 승격. arbiter 보상 가중치 수치는 이관하지 않는다.
6.2 코퍼스와 인프라
코퍼스 v1 목표 300+(ASSUMED, arbiter 240/300 미달 전례를 패밀리 생성기 선행으로 방지), held-out sha256 버킷 ~20% 동결. 커리큘럼은 WebRL 검증 패턴(실패-시드, critic 성공률 [0.05, 0.75] 밴드, MEASURED — 4.8%→42.4%) + R-Zero식 anti-repetition.
인프라: slime coding_agent_rl 구조 이식(별도 클린 샌드박스 채점, message tree fan-out), 훈련 경로는 verl(LoRA 일급) — TRL 철회·ROLL 폴백 결정 계승. verl GRPO+LoRA 불안정(#3226/#3784)의 재현-또는-해소는 착수 전 번다운. 비용 앵커: MiniMax-M1 풀 RL $534,700(MEASURED) 대비 본 프로그램은 두 자릿수 이상 작다. EI 1반복 ≈ 4.2k 에피소드, concurrency 12 기준 12–35시간(DERIVED, arbiter 실측 산식 재사용). rollout 비용이 예산을 지배 — V-BUD "cheap-rollout" 조건의 자체 충족 근거. 운영 규율: 한 번에 한 arm 평가(eval lock), 서빙 API shape 계약 테스트(CG22 전례), ABORTED 샘플 매니페스트 기록.
6.3 단계 로드맵
- Stage 0 — GEPA/프롬프트-최적화 대조군(스킵 불가): NPO "no universal winner" + mmGRPO 합성 우세(MEASURED, V-BUD). Exit — 코퍼스 동결+fingerprint 배선, base/GEPA-opt 비용 조정 곡선(teacher-LM 호출 포함 총예산 회계), capability-bound 판정: GEPA로 닫히지 않는 격차가 식별되면 진행, ~10% 이내로 닫히면 RL 중단.
- Stage 1 — SFT 콜드스타트(포맷): 능력이 아니라 레퍼토리 주입(스키마 준수, 복구 동사, 검증 노드 배치). Exit — 컴파일 통과율 ≥95%(ASSUMED, EI 선별 가능 조건에서 도출), 전 rl_viable 패밀리 pass@k > 0, provenance 테스트 green, full-SFT 과용 신호 없음.
- Stage 2 — Expert Iteration: arbiter 골격 재사용(n=12, gate-passer top-k, all-fail→hard_set, replay 25%). Exit — 목표군 ≥2 개선 + validity 유지(헤드라인 스칼라 아님), 플래토+잔여 분산, 해킹 시그니처 0건.
- Stage 3 — 온라인 GRPO(조건부): verl AgentLoop 네이티브, clip-higher, λ 실제 anneal, GiGPO+telescoping credit, 자동 halt 트립와이어. Exit — ≥200 안정 스텝, 사전등록 2x2 실행·보고(전이 실패도 유효한 결과 — 그 경우 산출물은 "DAG 도메인 자체 성능 + 전이 부재의 통제된 증거"로 재정의), Stage0 대비 파레토 지배점 부재 시 RL+prompt-opt 합성으로 회귀.
6.4 측정 계획
비용 조정 파레토 의무(V-EFF), pass^4 신뢰도(gpt-4o pass^1 ~61%→pass^8 <25% 붕괴 전례, MEASURED), 사전등록(2x2, scaffold ladder, random-reward+non-Qwen, utilization vector), 자동 실패 귀속은 훈련 신호 금지(step 14.2%). 통계 플로어: held-out 패밀리당 ~20태스크는 ~15pt 스윙만 탐지(MEASURED) — 최종 주장은 ≥3 seed + paired bootstrap. 평가기 fingerprint 핀, judge-vs-rubric 발산은 해킹 시그니처 취급.
7. 리스크·반론 (디베이트 verdict 요약)
- 디베이트의 4개 공격군(effort confound, construct mismatch, 예산/옵티마이저 선택, prior 재분배)은 어느 것도 기각되지 않았다 — 전부 "설계 제약으로 수용"이며, 그 제약의 실행 계획화가 곧 이 보고서의 로드맵이다.
- V-EFF·V-CON UPHELD: 비용 조정 파레토 없는 성과 주장과, "인터페이스 인접 전이 + 제약준수 규율"을 넘는 전이 서술은 이 프로그램에서 금지된다.
- V-BUD PARTIALLY ANSWERED: Stage0 GEPA arm 의무, RL 투자는 capability-bound 판정 + 저비용 rollout 경제 확인 후 — 단 가설 검증 목적의 RL은 가중치 업데이트 경로에서만 성립하므로 대체 불가다.
- V-CA2·V-CA3 제약 확정: 오케스트레이션은 ~8–10 엔티티/32K/39% 핸드오프 세율 임계 위에서만 정당하고, base가 못 보이는 행동은 SFT로 먼저 심으며, non-Qwen·random-reward·base 층화 없는 전이 주장은 무효다.
- 최상위 실행 리스크는 보상 해킹(R1)·effort confound(R2)·전이 실패(R3)이며, 전이 실패조차 통제된 부재 증거로서 독립 가치를 갖는다 — V-STOP의 stop-specific 보상 셀은 시간부패가 빠른 1순위 novelty 기회로 Stage3에 최소 실험 편입한다.
7.1 FINAL VERDICTS 요약 (debate-log 2026-08-31T09:30)
| Verdict | 판정 | 프로그램 반영 |
|---|---|---|
| V-EFF (SA2) | ATTACK UPHELD, 설계 제약 수용 | 모든 성과 = 비용 조정 파레토(토큰/순차토큰/pass^k/$); K2.5 수치는 vendor+effort-uncontrolled 이중 태그 |
| V-CON (SA1/3/5/10/12) | ATTACK UPHELD | 전이 주장 상한 = "인터페이스 인접 전이+제약준수 규율"; 네이티브 채점기·blinded artifact eval·스캐폴드 제거 사다리 의무. 반증 여지: RLTR 74.59>65.30 |
| V-BUD (CA1/SA6) | PARTIALLY ANSWERED | Stage0 GEPA arm 의무; RL은 capability-bound + cheap-rollout 조건; FLOPs-matched 비교의 첫 실행자 기회 |
| V-STOP | NARROWED, HOLDS | "stop-specific reward/credit으로 훈련된 정지 정책 없음"으로 재서술(Maestro는 outcome-rewarded stop만 — 선인용); novelty #1, 시간부패 높음 |
| V-CA2 | PARTIALLY UPHELD → 파티션 규칙 | ~8–10 엔티티/32K/39% 임계 위 breadth-first만 코퍼스 편입(DERIVED) |
| V-CA3 | UPHELD AS CONSTRAINT | non-Qwen·random-reward 더미·base 층화 사전등록; base 미표출 행동은 SFT 선행(subflow 0/8) |
| V-K2.5 | CG1–CG5 유지 | "1000 서브에이전트" REFUTED 교정; PARL은 CG26으로 정밀화 |
7.2 리스크 레지스터 (환경 장에서 통합)
| # | 리스크 | 심각도 | 완화 |
|---|---|---|---|
| R1 | 보상 해킹(fan-out 복제, node-stuffing, 자기채점 게이트) | 높음 | 규칙-검증 게이트 + anti-hacking ledger + λ 실제 annealing + caught-fault rate 보상(CG20) |
| R2 | effort confound — 성과가 토큰 구매로 판명 | 높음 | 비용 조정 파레토 + telescoping credit + CriticalSteps 회계(V-EFF) |
| R3 | 전이 실패(가설의 본질적 위험) | 높음 | 사전등록 2x2 조기 판정; 실패 시에도 DAG 도메인 자체 개선은 독립 가치(V-CON) |
| R4 | 과잉 오케스트레이션의 일반 태스크 유출 | 중간 | serial-optimal 함정 태스크 + no-op non-inferiority + non-regression battery(SA13/14) |
| R5 | verl GRPO+LoRA 불안정 | 중간 | 착수 전 재현-또는-해소; 폴백 = EI 연장 + iterative DPO |
| R6 | 워커 수준 성향은 오케스트레이터 훈련이 못 고침(FRT) | 중간 | frozen executor 체크포인트 선정의 평가 항목화; 워커 훈련은 스코프 아웃 |
| R7 | 판정자/시뮬레이터 파라미터 표류(temp 0.3 vs 0.7 전례) | 낮음 | fingerprint 포함, doc-code 단일 소스화 |
| R8 | stop-specific 셀 선점 경쟁(시간부패) | 중간 | Stage3에 stop-action + cost-adjusted return 최소 실험 편입(V-STOP) |
| R9 | 코퍼스 목표 미달 반복(240/300 전례) | 낮음 | 패밀리 생성기 선행, 생성기 없는 패밀리 v1 제외 |
8. 결론
가설은 살아 있고, 조건은 명시됐고, 도구는 착지해 있다. 종합 판정은 세 문장으로 압축된다.
- 주장할 수 있는 것: DAG-as-code 저작을 보상하는 RL은 위임 역학을 실제로 바꾸며(T1, 독립 재현), 인터페이스 인접 전이와 제약준수 규율까지는 현 증거가 지지한다. 그 이상 — task() 인터페이스로의 교차 전이(가설 A), 일반 플래닝 능력의 향상(가설 B) — 은 미증명이며, 이 보고서의 어떤 문장도 그 경계를 넘지 않는다.
- 해야 하는 것: Stage0 GEPA 대조군 → SFT 콜드스타트 → EI → 조건부 GRPO의 게이트식 로드맵을, 사전등록 2x2·스캐폴드 제거 사다리·non-Qwen+random-reward 통제·비용 조정 파레토·pass^4 아래에서 실행한다. 선행 조치 3건(보안 조치, API shape 정합, judge 배선)이 첫 스코어드 런의 전제다.
- 얻는 것: 어느 방향의 결과든 빈 연구 슬롯(CG6)의 첫 통제 증거이고, stop-specific 보상 셀(V-STOP)은 검증과 novelty 선점을 한 프로그램으로 겸하게 한다. 전이가 실패해도 산출물은 "DAG 도메인 자체 성능 + 전이 부재의 통제된 증거"로 성립한다.
조정 로그
섹션 초안 간 긴장·중복·표현 차이를 debate-log FINAL VERDICTS와 claim-graph 판정 기준으로 조정한 기록이다.
- 첫 섹션 명칭: 태스크 지시문의 "배경·남부현황"은 오타로 판단, 초안(arch-background.md) 원제의 "내부 상태/내부 자산"에 따라 배경·내부현황으로 확정했다.
- PlanningBench +18.01의 서술 수위: arch-hyp-a·arch-hyp-b·arch-reward가 각각 다른 강도로 인용 — claim-graph Gate pass 2의 CG8 REVISED(SUPPORTED→PARTIAL)에 따라 전 섹션에서 "objective isomorphism, 제약준수 규율 전이로 축소 서술"로 통일하고, 유효 잔존 발견은 결정성 ablation(+7.06 vs +0.75)만으로 한정했다.
- "1000 서브에이전트": 배경 초안은 회상 인용을 포함했으나 CG2 REFUTED에 따라 전 섹션 공통의 교정문(실제 상한 K2.5 100/1,500, K2.6 300/4,000; "1000"은 K2.6 코딩 데모 툴콜 수)으로 일원화하고 주장으로는 어디에서도 사용하지 않았다.
- Kimi 성과 수치의 태그: arch-cases의 WideSearch 72.7/72.8 내부 불일치는 초안의 결정(Table 6 값 72.7 채택)을 따랐고, 모든 스웜 수치에 CG4 PARTIAL의 vendor + effort-uncontrolled 이중 태그를 유지했다(V-EFF).
- FRT "프롬프트 무효" vs MAST "+15.6% 개입 효과"의 겉보기 모순: arch-cases의 해소(축이 다름 — 스캐폴드 구조는 유효한 레버, 프롬프트 문구는 죽은 레버; DERIVED, w2-agg C6)를 채택해 리스크 섹션과 일관시켰다.
- 가설 A/B의 판정 문구: 두 초안 모두 "promising, unproven"이나 승격 조건 표현이 달랐다 — debate-log 종합("4개 공격군이 요구하는 통제를 통과해야 주장 가능")에 맞춰 A는 3조건(2x2/사다리/prior 통제), B는 약한 버전 지지·강한 버전 미검증의 2단 판정으로 병렬 정리했다.
- 개입 횟수 신호: 배경 초안은 "D5 인접"으로 매핑 가능성을 남겼으나, arch-reward의 정밀 분석(은폐 유인, K2 F.3 과신 부작용과의 합성)에 따라 "훈련 보상에서 제외, 원인 페널티로 대체"를 최종 입장으로 채택했다.
- 보상 빈 셀의 개수: 배경 초안의 "빈 셀은 쌍(복구+중단)" 프레임을 유지하되, 복구는 보상 섹션의 결론(telescoping credit으로 credit 레이어에 흡수 — 보상함수가 아님)으로, 중단은 V-STOP의 novelty 슬롯으로 각각 귀속시켜 중복 서술을 제거했다.
- FlowReasoner: CG29 REFUTED(가중치 상수 0.7/0.15/0.15 NOT-FOUND-IN-SOURCE)에 따라 어떤 섹션에서도 인용하지 않았다. arch-hyp-b의 do-not-cite 언급만 인용 위생 규칙으로 반영했다.
- annex/UNRESOLVED 항목: $4.6M 학습비(CG14), Researcher=K2 Thinking 전신(CG15), CG18 미검증군, 정크 저널 2건은 "인용하지 않음" 또는 "단정하지 않음"으로만 언급하고 수치·주장으로 사용하지 않았다.
- ASSUMED 값의 유지: 코퍼스 구성 비율(결정적 최적해 ~50%, serial-optimal ~15%), SFT exit 기준(컴파일 ≥95%), credit 작동점(GiGPO ω≈0.8, TRACE K=3/γ=0.8)은 초안의 ASSUMED 태그와 도출 근거를 그대로 보존했다 — 첫 스윕 대상임을 명시한다.
부록 A. 방법론 (자동 생성 초안)
수집·검증 파이프라인
본 보고서는 mass ulw-research 프로토콜(팀 기반 최대 포화 리서치)로 수집되었다.
- Wave 1 (DAG
dag_126d8983): 49개 노드 — 내부 코드베이스 14개(rl-env 설계문서·코드, omo workflow 툴, arbiter-pi 스킬, storm/carrier 스펙, 스레드 다이제스트) + 외부 리서치 30개(Kimi 클러스터 6, 타 랩 4, RL 방법론·보상·전이·플래닝·벤치마크 20) + 집계 5. 카테고리: quick/unspecified-low/unspecified-high/deep 혼합. - Wave 2 (DAG
dag_3692d153): 17개 노드 — PARL 논문 정밀 추출, OpenPARL 재현, 오케스트레이션-RL 서베이(2605.02801), MAST, Anthropic Red Team, RLTR, PlanningBench 정독, matched-budget 반례 검색, stopping-decision 공백 검증, 유틸리티 지표 카탈로그 등 + 집계. - 복구/재검증: 노드 불량(w2-kimi-k26-k3) 복구 레인, recall-only 수치(FlowReasoner/AFM) 재검증 레인. 회상된 상수 1건 REFUTED 처리.
- 디베이트 팀 (29ceb433): skeptic(ultrabrain) + contrarian(deep) — 가설 A/B·보상설계·프레이밍을 각 1라운드 이상 공격, verdict를 debate-log.md에 기록.
- 클레임 게이트 (claim-graph.md): 고위험 주장 34건 심사 — 지지(verified-claims) / 부분 / 부록 annex(unresolved·refuted). "1000 서브에이전트" 오인용 정정 등.
- 출처: 정제 후 760개 인용-grade URL (76개 고유 도메인). 각 수치는 MEASURED/ASSUMED/DERIVED 리니지 태그.
신뢰도 표기
- vendor-self-reported: 벤더 자체 보고 수치(독립 재현 없음)
- 조건 태그: GAIA 등 조건 상이 수치는 평가 조건 명시 필수
- do-not-cite 목록: junk-venue 논문 2종, 미확인 수치 4건 인용 금지
- 검증되지 않은 주장은 본문이 아닌 미결 annex에만 기록
노드 실패 기록 (투명성)
- arch-hyp-a/b/reward/env/cases 중 4건이 provider stream timeout으로 실패했으나, 실패 전 산출물(섹션 초안)은 디스크에 완전히 기록되어 있어 amend로 reducer만 재실행. 초안 완전성은 오케스트레이터가 직접 파일을 읽어 검증.
출처
출처 전체 보기 (760개)
| 번호 | URL |
|---|---|
| S1 | http://export.arxiv.org/api/query — in: w2-orch-survey |
| S2 | http://export.arxiv.org/api/query?id_list=2605.20873 — in: w2-planningbench |
| S3 | http://export.arxiv.org/api/query?search_query=all:%22Kimi%20K2%22 — in: ext-kimi-k2 |
| S4 | http://web.archive.org/cdx/search/cdx?url=openpipe.ai/blog* — in: ext-llm-judge-rubric |
| S5 | http://web.archive.org/web/20250621104403/https://moonshotai.github.io/Kimi-Researcher/ — in: ext-kimi-researcher |
| S6 | http://web.archive.org/web/20260813221311/https://moonshotai.github.io/Kimi-K2/thinking.html — in: w2-kimi-k26-k3 |
| S7 | https://aclanthology.org/2026.acl-long.1962/ — in: ext-credit-assignment |
| S8 | https://api-docs.deepseek.com/news/news250821 — in: ext-glm-qwen |
| S9 | https://api-docs.deepseek.com/news/news250929 — in: ext-glm-qwen |
| S10 | https://api.crossref.org — in: w2-stopping |
| S11 | https://api.crossref.org/works — in: ext-workflow-search,ext-planning-training |
| S12 | https://api.crossref.org/works?query=PlanningBench+planning — in: w2-planningbench |
| S13 | https://api.crossref.org/works/10.71465/ajbd3639 — in: w2-crosstask |
| S14 | https://api.crossref.org/works/10.71465/ajdsa3667 — in: w2-crosstask |
| S15 | https://api.github.com/repos/Ayanami0730/deep_research_bench/readme — in: ext-benchmarks |
| S16 | https://api.github.com/repos/GuanxingLu/miles/commits?sha=openparl-v1 — in: w2-openparl |
| S17 | https://api.github.com/repos/GuanxingLu/OpenPARL — in: w2-openparl |
| S18 | https://api.github.com/repos/microsoft/agent-lightning/git/trees/v0.x — in: ext-frameworks |
| S19 | https://api.github.com/search/repositories — in: ext-workflow-search |
| S20 | https://api.github.com/search/repositories?q=PlanBench — in: ext-planning-training |
| S21 | https://api.observablehq.com/@tomlarkworthy/gepa.js — in: w2-matched-budget |
| S22 | https://api.openalex.org/sources/S5407050933 — in: w2-crosstask |
| S23 | https://api.openalex.org/sources/S5407055466 — in: w2-crosstask |
| S24 | https://api.openalex.org/works?search=GEPA+GRPO&filter=from_publication_date:2025-09-01 — in: w2-matched-budget |
| S25 | https://api.openalex.org/works?search=PlanningBench%20verifiable%20planning — in: w2-planningbench |
| S26 | https://api.openalex.org/works/W4414971614 — in: w2-matched-budget |
| S27 | https://api.openalex.org/works/W4416711932 — in: w2-rltr |
| S28 | https://api.openalex.org/works/W7078199011 — in: w2-rltr |
| S29 | https://api.openalex.org/works/W7139106429 — in: w2-crosstask |
| S30 | https://api.openalex.org/works/W7154311166 — in: w2-crosstask |
| S31 | https://api.openalex.org/works/W7162092528 — in: w2-matched-budget |
| S32 | https://api.openalex.org/works/W7162893618 — in: w2-matched-budget |
| S33 | https://api.openalex.org/works/W7163597140 — in: w2-matched-budget |
| S34 | https://api.semanticscholar.org/graph/v1/author/search?query=David+Duvenaud — in: w2-crosstask |
| S35 | https://api.semanticscholar.org/graph/v1/paper/arXiv:2411.02337/citations — in: ext-webrl |
| S36 | https://api.semanticscholar.org/graph/v1/paper/arXiv:2503.13657 — in: w2-mast |
| S37 | https://api.semanticscholar.org/graph/v1/paper/arXiv:2508.19598/citations — in: w2-rltr |
| S38 | https://api.semanticscholar.org/graph/v1/paper/arXiv:2508.19598/references — in: w2-rltr |
| S39 | https://api.semanticscholar.org/graph/v1/paper/arXiv:2509.02547/citations — in: w2-crosstask |
| S40 | https://api.semanticscholar.org/graph/v1/paper/arXiv:2605.20873 — in: w2-planningbench |
| S41 | https://api.semanticscholar.org/graph/v1/paper/DOI:10.71465/ajbd3639 — in: w2-crosstask |
| S42 | https://api.semanticscholar.org/graph/v1/paper/DOI:10.71465/ajdsa3667 — in: w2-crosstask |
| S43 | https://api.semanticscholar.org/graph/v1/paper/search — in: ext-workflow-search |
| S44 | https://ar5iv.labs.arxiv.org/html/2209.07753 — in: ext-codeact |
| S45 | https://ar5iv.labs.arxiv.org/html/2210.03629 — in: ext-codeact |
| S46 | https://ar5iv.labs.arxiv.org/html/2211.10435 — in: ext-codeact |
| S47 | https://ar5iv.labs.arxiv.org/html/2211.12588 — in: ext-codeact |
| S48 | https://ar5iv.labs.arxiv.org/html/2305.04091 — in: ext-codeact |
| S49 | https://ar5iv.labs.arxiv.org/html/2305.16291 — in: ext-codeact |
| S50 | https://ar5iv.labs.arxiv.org/html/2309.12499 — in: ext-codeact |
| S51 | https://ar5iv.labs.arxiv.org/html/2310.02170 — in: w2-util-metrics |
| S52 | https://ar5iv.labs.arxiv.org/html/2312.04511 — in: w2-util-metrics |
| S53 | https://ar5iv.labs.arxiv.org/html/2402.01030 — in: ext-codeact |
| S54 | https://ar5iv.labs.arxiv.org/html/2402.05120 — in: w2-util-metrics |
| S55 | https://ar5iv.labs.arxiv.org/html/2402.16823 — in: w2-util-metrics,ext-multiagent-rl |
| S56 | https://ar5iv.labs.arxiv.org/html/2403.12031 — in: w2-util-metrics |
| S57 | https://ar5iv.labs.arxiv.org/html/2406.07155 — in: w2-util-metrics |
| S58 | https://ar5iv.labs.arxiv.org/html/2406.12045 — in: w2-util-metrics |
| S59 | https://ar5iv.labs.arxiv.org/html/2406.18665 — in: w2-util-metrics |
| S60 | https://ar5iv.labs.arxiv.org/html/2407.01502 — in: w2-util-metrics |
| S61 | https://ar5iv.labs.arxiv.org/html/2410.07095 — in: w2-util-metrics |
| S62 | https://ar5iv.labs.arxiv.org/html/2410.10762 — in: w2-util-metrics |
| S63 | https://ar5iv.labs.arxiv.org/html/2412.01928 — in: ext-multiagent-rl |
| S64 | https://ar5iv.labs.arxiv.org/html/2501.17161 — in: ext-transfer-evidence |
| S65 | https://ar5iv.labs.arxiv.org/html/2502.11133 — in: ext-multiagent-rl |
| S66 | https://ar5iv.labs.arxiv.org/html/2502.18439 — in: ext-multiagent-rl |
| S67 | https://ar5iv.labs.arxiv.org/html/2502.18449 — in: ext-transfer-evidence,ext-swe-rl |
| S68 | https://ar5iv.labs.arxiv.org/html/2502.19411 — in: ext-codeact |
| S69 | https://ar5iv.labs.arxiv.org/html/2503.05592 — in: ext-transfer-evidence |
| S70 | https://ar5iv.labs.arxiv.org/html/2503.09516 — in: ext-transfer-evidence |
| S71 | https://ar5iv.labs.arxiv.org/html/2503.13657 — in: w2-util-metrics,w2-mast |
| S72 | https://ar5iv.labs.arxiv.org/html/2503.23829 — in: ext-transfer-evidence |
| S73 | https://ar5iv.labs.arxiv.org/html/2504.07912 — in: ext-transfer-evidence |
| S74 | https://ar5iv.labs.arxiv.org/html/2504.13837 — in: ext-transfer-evidence |
| S75 | https://ar5iv.labs.arxiv.org/html/2504.15257 — in: ext-multiagent-rl |
| S76 | https://ar5iv.labs.arxiv.org/html/2505.00212 — in: w2-util-metrics |
| S77 | https://ar5iv.labs.arxiv.org/html/2505.14652 — in: ext-transfer-evidence |
| S78 | https://ar5iv.labs.arxiv.org/html/2505.18129 — in: ext-transfer-evidence |
| S79 | https://ar5iv.labs.arxiv.org/html/2505.19591 — in: ext-multiagent-rl |
| S80 | https://ar5iv.labs.arxiv.org/html/2505.22617 — in: ext-transfer-evidence |
| S81 | https://ar5iv.labs.arxiv.org/html/2505.24864 — in: ext-transfer-evidence |
| S82 | https://ar5iv.labs.arxiv.org/html/2506.07982 — in: w2-util-metrics |
| S83 | https://ar5iv.labs.arxiv.org/html/2506.14245 — in: ext-transfer-evidence |
| S84 | https://ar5iv.labs.arxiv.org/html/2507.02592 — in: ext-search-rl |
| S85 | https://ar5iv.labs.arxiv.org/html/2508.13167 — in: ext-multiagent-rl |
| S86 | https://ar5iv.labs.arxiv.org/html/2509.04259 — in: ext-transfer-evidence |
| S87 | https://ar5iv.labs.arxiv.org/html/2509.05368 — in: ext-codeact |
| S88 | https://arctic-shift.com/api/posts/search — in: ext-kimi-community |
| S89 | https://arstechnica.com/ai/2025/01/openai-launches-operator-an-ai-agent-that-can-operate-your-computer/ — in: ext-openai-agents |
| S90 | https://artificialanalysis.ai/models/kimi-k2-5 — in: ext-kimi-community |
| S91 | https://arxiv.org/abs/1603.08983 — in: w2-stopping |
| S92 | https://arxiv.org/abs/1609.05140 — in: w2-stopping |
| S93 | https://arxiv.org/abs/1706.06195 — in: w2-stopping |
| S94 | https://arxiv.org/abs/2107.05407 — in: w2-stopping |
| S95 | https://arxiv.org/abs/2204.01691 — in: ext-planning-training |
| S96 | https://arxiv.org/abs/2206.10498 — in: ext-benchmarks,ext-planning-training |
| S97 | https://arxiv.org/abs/2209.07753 — in: ext-codeact |
| S98 | https://arxiv.org/abs/2210.03629 — in: ext-codeact |
| S99 | https://arxiv.org/abs/2210.10765 — in: w2-stopping |
| S100 | https://arxiv.org/abs/2211.10435 — in: ext-codeact |
| S101 | https://arxiv.org/abs/2211.11603 — in: ext-credit-assignment |
| S102 | https://arxiv.org/abs/2211.12588 — in: ext-codeact |
| S103 | https://arxiv.org/abs/2212.08073 — in: ext-llm-judge-rubric |
| S104 | https://arxiv.org/abs/2302.06706 — in: ext-benchmarks |
| S105 | https://arxiv.org/abs/2303.16634 — in: ext-llm-judge-rubric |
| S106 | https://arxiv.org/abs/2304.09870 — in: ext-multiagent-rl |
| S107 | https://arxiv.org/abs/2304.11477 — in: ext-planning-training |
| S108 | https://arxiv.org/abs/2305.04091 — in: ext-codeact |
| S109 | https://arxiv.org/abs/2305.16291 — in: ext-codeact |
| S110 | https://arxiv.org/abs/2305.16653 — in: ext-planning-training |
| S111 | https://arxiv.org/abs/2305.17926 — in: ext-llm-judge-rubric |
| S112 | https://arxiv.org/abs/2305.20050 — in: ext-prm-vs-orm,ext-llm-judge-rubric |
| S113 | https://arxiv.org/abs/2306.05685 — in: ext-llm-judge-rubric |
| S114 | https://arxiv.org/abs/2307.13854 — in: ext-webrl |
| S115 | https://arxiv.org/abs/2308.15452 — in: ext-codeact |
| S116 | https://arxiv.org/abs/2309.12499 — in: ext-codeact |
| S117 | https://arxiv.org/abs/2310.03714 — in: ext-workflow-search |
| S118 | https://arxiv.org/abs/2310.05915 — in: ext-planning-training |
| S119 | https://arxiv.org/abs/2310.06770 — in: ext-benchmarks |
| S120 | https://arxiv.org/abs/2310.08118 — in: ext-planning-training |
| S121 | https://arxiv.org/abs/2310.12823 — in: ext-planning-training |
| S122 | https://arxiv.org/abs/2311.12983 — in: ext-benchmarks |
| S123 | https://arxiv.org/abs/2312.08935 — in: ext-prm-vs-orm |
| S124 | https://arxiv.org/abs/2401.10020 — in: ext-llm-judge-rubric |
| S125 | https://arxiv.org/abs/2402.01030 — in: ext-codeact |
| S126 | https://arxiv.org/abs/2402.01622 — in: ext-planning-training |
| S127 | https://arxiv.org/abs/2402.01817 — in: ext-planning-training |
| S128 | https://arxiv.org/abs/2402.05120 — in: ext-cognition-counter |
| S129 | https://arxiv.org/abs/2402.16823 — in: ext-multiagent-rl,ext-workflow-search,ext-planning-training |
| S130 | https://arxiv.org/abs/2402.19446 — in: w2-stopping,ext-prm-vs-orm,ext-credit-assignment,ext-planning-training |
| S131 | https://arxiv.org/abs/2404.03648 — in: ext-webrl |
| S132 | https://arxiv.org/abs/2404.04475 — in: ext-llm-judge-rubric |
| S133 | https://arxiv.org/abs/2404.06654 — in: w2-context-bounds |
| S134 | https://arxiv.org/abs/2404.13076 — in: ext-llm-judge-rubric |
| S135 | https://arxiv.org/abs/2405.01535 — in: ext-llm-judge-rubric |
| S136 | https://arxiv.org/abs/2405.04215 — in: ext-planning-training |
| S137 | https://arxiv.org/abs/2406.02818 — in: ext-multiagent-rl |
| S138 | https://arxiv.org/abs/2406.04151 — in: ext-planning-training |
| S139 | https://arxiv.org/abs/2406.04520 — in: ext-benchmarks,ext-planning-training |
| S140 | https://arxiv.org/abs/2406.06592 — in: ext-prm-vs-orm |
| S141 | https://arxiv.org/abs/2406.11695 — in: ext-workflow-search |
| S142 | https://arxiv.org/abs/2406.12045 — in: ext-benchmarks |
| S143 | https://arxiv.org/abs/2407.01489 — in: w2-mast |
| S144 | https://arxiv.org/abs/2407.01502 — in: w2-mast |
| S145 | https://arxiv.org/abs/2407.10817 — in: ext-llm-judge-rubric |
| S146 | https://arxiv.org/abs/2407.12036 — in: ext-codeact |
| S147 | https://arxiv.org/abs/2407.16741 — in: ext-swe-rl |
| S148 | https://arxiv.org/abs/2408.00764 — in: ext-webrl |
| S149 | https://arxiv.org/abs/2408.07199 — in: ext-prm-vs-orm,ext-planning-training |
| S150 | https://arxiv.org/abs/2408.08435 — in: ext-multiagent-rl,ext-workflow-search,ext-planning-training |
| S151 | https://arxiv.org/abs/2408.15240 — in: ext-llm-judge-rubric |
| S152 | https://arxiv.org/abs/2409.00920 — in: ext-toolrl |
| S153 | https://arxiv.org/abs/2409.13373 — in: ext-planning-training |
| S154 | https://arxiv.org/abs/2409.19256 — in: ext-frameworks,w2-slime-verl |
| S155 | https://arxiv.org/abs/2410.07095 — in: ext-benchmarks |
| S156 | https://arxiv.org/abs/2410.08115 — in: ext-multiagent-rl |
| S157 | https://arxiv.org/abs/2410.08146 — in: ext-prm-vs-orm |
| S158 | https://arxiv.org/abs/2410.10762 — in: w2-matched-budget,ext-multiagent-rl,ext-workflow-search,ext-planning-training |
| S159 | https://arxiv.org/abs/2410.11782 — in: ext-multiagent-rl |
| S160 | https://arxiv.org/abs/2410.12784 — in: ext-llm-judge-rubric |
| S161 | https://arxiv.org/abs/2411.00820 — in: ext-webrl |
| S162 | https://arxiv.org/abs/2411.02337 — in: ext-webrl,ext-planning-training |
| S163 | https://arxiv.org/abs/2411.14503 — in: ext-codeact |
| S164 | https://arxiv.org/abs/2411.15124 — in: ext-rlvr-tulu |
| S165 | https://arxiv.org/abs/2412.01928 — in: ext-multiagent-rl |
| S166 | https://arxiv.org/abs/2412.01981 — in: ext-prm-vs-orm |
| S167 | https://arxiv.org/abs/2412.06559 — in: ext-prm-vs-orm |
| S168 | https://arxiv.org/abs/2412.11605 — in: ext-curriculum-taskgen |
| S169 | https://arxiv.org/abs/2412.13682 — in: w2-planningbench |
| S170 | https://arxiv.org/abs/2412.21139 — in: ext-swe-rl |
| S171 | https://arxiv.org/abs/2501.07301 — in: ext-prm-vs-orm |
| S172 | https://arxiv.org/abs/2501.07834 — in: ext-codeact |
| S173 | https://arxiv.org/abs/2501.11425 — in: ext-credit-assignment |
| S174 | https://arxiv.org/abs/2501.12599 — in: ext-kimi-k15,ext-llm-judge-rubric,ext-kimi-k2 |
| S175 | https://arxiv.org/abs/2501.12948 — in: ext-rlvr-tulu,ext-prm-vs-orm,ext-llm-judge-rubric,ext-glm-qwen |
| S176 | https://arxiv.org/abs/2501.17161 — in: ext-transfer-evidence,ext-planning-training |
| S177 | https://arxiv.org/abs/2501.17167 — in: ext-codeact |
| S178 | https://arxiv.org/abs/2502.01456 — in: ext-credit-assignment |
| S179 | https://arxiv.org/abs/2502.01600 — in: ext-credit-assignment |
| S180 | https://arxiv.org/abs/2502.04180 — in: ext-multiagent-rl,ext-workflow-search |
| S181 | https://arxiv.org/abs/2502.04306 — in: ext-multiagent-rl |
| S182 | https://arxiv.org/abs/2502.05167 — in: w2-context-bounds |
| S183 | https://arxiv.org/abs/2502.05664 — in: ext-codeact |
| S184 | https://arxiv.org/abs/2502.08235 — in: w2-stopping |
| S185 | https://arxiv.org/abs/2502.10325 — in: ext-prm-vs-orm,w2-rltr,ext-credit-assignment |
| S186 | https://arxiv.org/abs/2502.11133 — in: ext-multiagent-rl |
| S187 | https://arxiv.org/abs/2502.11221 — in: ext-planning-training |
| S188 | https://arxiv.org/abs/2502.16111 — in: ext-planning-training |
| S189 | https://arxiv.org/abs/2502.18439 — in: ext-multiagent-rl |
| S190 | https://arxiv.org/abs/2502.18449 — in: ext-transfer-evidence,ext-swe-rl |
| S191 | https://arxiv.org/abs/2502.19411 — in: ext-codeact |
| S192 | https://arxiv.org/abs/2503.02682 — in: ext-planning-training |
| S193 | https://arxiv.org/abs/2503.03686 — in: ext-multiagent-rl |
| S194 | https://arxiv.org/abs/2503.04697 — in: w2-stopping |
| S195 | https://arxiv.org/abs/2503.05592 — in: ext-transfer-evidence |
| S196 | https://arxiv.org/abs/2503.09501 — in: ext-multiagent-rl |
| S197 | https://arxiv.org/abs/2503.09516 — in: ext-transfer-evidence,w2-stopping |
| S198 | https://arxiv.org/abs/2503.09572 — in: ext-planning-training |
| S199 | https://arxiv.org/abs/2503.13657 — in: ext-cognition-counter,w2-mast |
| S200 | https://arxiv.org/abs/2503.14476 — in: ext-transfer-evidence,ext-kimi-k15,ext-credit-assignment |
| S201 | https://arxiv.org/abs/2503.15478 — in: ext-prm-vs-orm,ext-credit-assignment |
| S202 | https://arxiv.org/abs/2503.16419 — in: w2-stopping |
| S203 | https://arxiv.org/abs/2503.19470 — in: ext-search-rl |
| S204 | https://arxiv.org/abs/2503.20783 — in: ext-transfer-evidence,ext-rlvr-tulu,ext-kimi-k15,ext-credit-assignment |
| S205 | https://arxiv.org/abs/2503.23383 — in: ext-transfer-evidence,ext-toolrl |
| S206 | https://arxiv.org/abs/2503.23829 — in: ext-transfer-evidence |
| S207 | https://arxiv.org/abs/2504.01400 — in: ext-toolrl |
| S208 | https://arxiv.org/abs/2504.04736 — in: ext-planning-training |
| S209 | https://arxiv.org/abs/2504.07164 — in: ext-swe-rl,w2-slime-verl |
| S210 | https://arxiv.org/abs/2504.07912 — in: ext-transfer-evidence |
| S211 | https://arxiv.org/abs/2504.10449 — in: ext-multiagent-rl |
| S212 | https://arxiv.org/abs/2504.11536 — in: ext-transfer-evidence,ext-toolrl |
| S213 | https://arxiv.org/abs/2504.12516 — in: ext-benchmarks |
| S214 | https://arxiv.org/abs/2504.13171 — in: w2-context-bounds |
| S215 | https://arxiv.org/abs/2504.13837 — in: ext-transfer-evidence,ext-rlvr-tulu,ext-planning-training |
| S216 | https://arxiv.org/abs/2504.13958 — in: ext-transfer-evidence,ext-agentic-rl-survey,w2-rltr,ext-toolrl |
| S217 | https://arxiv.org/abs/2504.14773 — in: ext-planning-training |
| S218 | https://arxiv.org/abs/2504.14870 — in: w2-rltr |
| S219 | https://arxiv.org/abs/2504.15257 — in: ext-multiagent-rl,ext-workflow-search,w2-flowreasoner-afm-verify |
| S220 | https://arxiv.org/abs/2504.15466 — in: ext-parallel-thinking |
| S221 | https://arxiv.org/abs/2504.15895 — in: w2-stopping |
| S222 | https://arxiv.org/abs/2504.20073 — in: w2-stopping,ext-frameworks,ext-prm-vs-orm,ext-credit-assignment |
| S223 | https://arxiv.org/abs/2504.20571 — in: ext-transfer-evidence |
| S224 | https://arxiv.org/abs/2504.21798 — in: ext-swe-rl |
| S225 | https://arxiv.org/abs/2505.00024 — in: ext-toolrl |
| S226 | https://arxiv.org/abs/2505.01441 — in: ext-toolrl |
| S227 | https://arxiv.org/abs/2505.01479 — in: ext-planning-training |
| S228 | https://arxiv.org/abs/2505.04588 — in: ext-transfer-evidence,ext-search-rl |
| S229 | https://arxiv.org/abs/2505.06120 — in: w2-context-bounds |
| S230 | https://arxiv.org/abs/2505.07512 — in: ext-toolrl |
| S231 | https://arxiv.org/abs/2505.07686 — in: w2-stopping |
| S232 | https://arxiv.org/abs/2505.09388 — in: ext-glm-qwen |
| S233 | https://arxiv.org/abs/2505.10475 — in: ext-multiagent-rl,ext-parallel-thinking |
| S234 | https://arxiv.org/abs/2505.10978 — in: w2-stopping,ext-agentic-rl-survey,ext-prm-vs-orm,ext-credit-assignment |
| S235 | https://arxiv.org/abs/2505.11821 — in: ext-credit-assignment |
| S236 | https://arxiv.org/abs/2505.14652 — in: ext-transfer-evidence |
| S237 | https://arxiv.org/abs/2505.15340 — in: ext-parallel-thinking |
| S238 | https://arxiv.org/abs/2505.16400 — in: ext-transfer-evidence |
| S239 | https://arxiv.org/abs/2505.16410 — in: ext-toolrl |
| S240 | https://arxiv.org/abs/2505.18129 — in: ext-transfer-evidence |
| S241 | https://arxiv.org/abs/2505.19591 — in: ext-multiagent-rl |
| S242 | https://arxiv.org/abs/2505.22617 — in: ext-transfer-evidence |
| S243 | https://arxiv.org/abs/2505.23564 — in: ext-credit-assignment |
| S244 | https://arxiv.org/abs/2505.24864 — in: ext-transfer-evidence |
| S245 | https://arxiv.org/abs/2506.01939 — in: ext-transfer-evidence |
| S246 | https://arxiv.org/abs/2506.03570 — in: ext-prm-vs-orm |
| S247 | https://arxiv.org/abs/2506.07982 — in: ext-benchmarks |
| S248 | https://arxiv.org/abs/2506.09991 — in: ext-multiagent-rl |
| S249 | https://arxiv.org/abs/2506.10947 — in: ext-transfer-evidence,ext-rlvr-tulu |
| S250 | https://arxiv.org/abs/2506.11763 — in: ext-benchmarks |
| S251 | https://arxiv.org/abs/2506.13585 — in: w2-slime-verl,ext-glm-qwen |
| S252 | https://arxiv.org/abs/2506.14245 — in: ext-transfer-evidence |
| S253 | https://arxiv.org/abs/2506.15672 — in: ext-multiagent-rl |
| S254 | https://arxiv.org/abs/2506.16507 — in: ext-llm-judge-rubric |
| S255 | https://arxiv.org/abs/2506.18254 — in: ext-transfer-evidence |
| S256 | https://arxiv.org/abs/2507.02592 — in: ext-webrl |
| S257 | https://arxiv.org/abs/2507.10532 — in: ext-rlvr-tulu |
| S258 | https://arxiv.org/abs/2507.17307 — in: ext-credit-assignment |
| S259 | https://arxiv.org/abs/2507.17746 — in: ext-llm-judge-rubric |
| S260 | https://arxiv.org/abs/2507.18071 — in: ext-credit-assignment |
| S261 | https://arxiv.org/abs/2507.19457 — in: w2-matched-budget,ext-workflow-search |
| S262 | https://arxiv.org/abs/2507.19849 — in: ext-credit-assignment |
| S263 | https://arxiv.org/abs/2507.20534 — in: ext-llm-judge-rubric,ext-kimi-k2 |
| S264 | https://arxiv.org/abs/2508.06471 — in: ext-glm-qwen |
| S265 | https://arxiv.org/abs/2508.07976 — in: ext-search-rl |
| S266 | https://arxiv.org/abs/2508.07999 — in: w2-openparl |
| S267 | https://arxiv.org/abs/2508.12685 — in: ext-toolrl |
| S268 | https://arxiv.org/abs/2508.13167 — in: ext-multiagent-rl,w2-flowreasoner-afm-verify |
| S269 | https://arxiv.org/abs/2508.19598 — in: ext-agentic-rl-survey,w2-rltr |
| S270 | https://arxiv.org/abs/2508.20404 — in: ext-frameworks |
| S271 | https://arxiv.org/abs/2509.02479 — in: ext-toolrl |
| S272 | https://arxiv.org/abs/2509.02547 — in: ext-agentic-rl-survey,w2-rltr,ext-credit-assignment |
| S273 | https://arxiv.org/abs/2509.04259 — in: ext-transfer-evidence |
| S274 | https://arxiv.org/abs/2509.04475 — in: ext-parallel-thinking |
| S275 | https://arxiv.org/abs/2509.04642 — in: w2-matched-budget |
| S276 | https://arxiv.org/abs/2509.05368 — in: ext-codeact |
| S277 | https://arxiv.org/abs/2509.06733 — in: ext-agentic-rl-survey |
| S278 | https://arxiv.org/abs/2509.07980 — in: ext-parallel-thinking |
| S279 | https://arxiv.org/abs/2509.08483 — in: ext-parallel-thinking |
| S280 | https://arxiv.org/abs/2509.08755 — in: ext-frameworks,ext-credit-assignment,ext-planning-training |
| S281 | https://arxiv.org/abs/2509.10550 — in: w2-stopping |
| S282 | https://arxiv.org/abs/2509.19199 — in: ext-credit-assignment |
| S283 | https://arxiv.org/abs/2509.20616 — in: ext-agentic-rl-survey |
| S284 | https://arxiv.org/abs/2509.21240 — in: ext-parallel-thinking,ext-credit-assignment |
| S285 | https://arxiv.org/abs/2509.25140 — in: ext-planning-training |
| S286 | https://arxiv.org/abs/2510.00219 — in: ext-parallel-thinking |
| S287 | https://arxiv.org/abs/2510.00263 — in: ext-llm-judge-rubric |
| S288 | https://arxiv.org/abs/2510.01394 — in: w2-stopping |
| S289 | https://arxiv.org/abs/2510.07743 — in: ext-llm-judge-rubric |
| S290 | https://arxiv.org/abs/2510.08049 — in: ext-prm-vs-orm |
| S291 | https://arxiv.org/abs/2510.13786 — in: ext-parallel-thinking |
| S292 | https://arxiv.org/abs/2510.15719 — in: w2-stopping |
| S293 | https://arxiv.org/abs/2510.16724 — in: ext-agentic-rl-survey |
| S294 | https://arxiv.org/abs/2510.17314 — in: ext-llm-judge-rubric |
| S295 | https://arxiv.org/abs/2510.24698 — in: ext-parallel-thinking |
| S296 | https://arxiv.org/abs/2511.01181 — in: w2-stopping |
| S297 | https://arxiv.org/abs/2511.08325 — in: ext-prm-vs-orm |
| S298 | https://arxiv.org/abs/2511.14846 — in: ext-credit-assignment |
| S299 | https://arxiv.org/abs/2511.16108 — in: ext-swe-rl,ext-frameworks |
| S300 | https://arxiv.org/abs/2512.02038 — in: ext-agentic-rl-survey |
| S301 | https://arxiv.org/abs/2512.07461 — in: ext-parallel-thinking |
| S302 | https://arxiv.org/abs/2512.07843 — in: ext-parallel-thinking |
| S303 | https://arxiv.org/abs/2512.17008 — in: ext-credit-assignment |
| S304 | https://arxiv.org/abs/2512.23707 — in: ext-llm-judge-rubric |
| S305 | https://arxiv.org/abs/2601.05593 — in: ext-parallel-thinking |
| S306 | https://arxiv.org/abs/2601.08654 — in: ext-llm-judge-rubric |
| S307 | https://arxiv.org/abs/2601.12538 — in: ext-agentic-rl-survey |
| S308 | https://arxiv.org/abs/2601.14652 — in: w2-stopping |
| S309 | https://arxiv.org/abs/2601.18137 — in: ext-planning-training |
| S310 | https://arxiv.org/abs/2601.21619 — in: ext-parallel-thinking |
| S311 | https://arxiv.org/abs/2602.02276 — in: w2-stopping,ext-kimi-k25,ext-kimi-swarm-blog,w2-openparl |
| S312 | https://arxiv.org/abs/2602.03845 — in: ext-parallel-thinking |
| S313 | https://arxiv.org/abs/2602.04634 — in: w2-openparl |
| S314 | https://arxiv.org/abs/2602.06795 — in: ext-llm-judge-rubric |
| S315 | https://arxiv.org/abs/2602.07839 — in: w2-rltr |
| S316 | https://arxiv.org/abs/2602.08344 — in: ext-parallel-thinking |
| S317 | https://arxiv.org/abs/2602.08847 — in: w2-stopping |
| S318 | https://arxiv.org/abs/2602.09514 — in: ext-planning-training |
| S319 | https://arxiv.org/abs/2602.11114 — in: ext-workflow-search |
| S320 | https://arxiv.org/abs/2602.11767 — in: ext-credit-assignment |
| S321 | https://arxiv.org/abs/2602.17547 — in: ext-credit-assignment |
| S322 | https://arxiv.org/abs/2603.01914 — in: w2-stopping |
| S323 | https://arxiv.org/abs/2603.06194 — in: ext-credit-assignment |
| S324 | https://arxiv.org/abs/2603.08754 — in: ext-credit-assignment |
| S325 | https://arxiv.org/abs/2603.19685 — in: ext-planning-training |
| S326 | https://arxiv.org/abs/2603.21972 — in: ext-prm-vs-orm |
| S327 | https://arxiv.org/abs/2604.01302 — in: ext-parallel-thinking |
| S328 | https://arxiv.org/abs/2604.02226 — in: w2-stopping |
| S329 | https://arxiv.org/abs/2604.09459 — in: ext-credit-assignment |
| S330 | https://arxiv.org/abs/2604.13618 — in: ext-llm-judge-rubric |
| S331 | https://arxiv.org/abs/2604.13946 — in: ext-codeact |
| S332 | https://arxiv.org/abs/2604.16029 — in: ext-parallel-thinking |
| S333 | https://arxiv.org/abs/2604.19756 — in: ext-workflow-search |
| S334 | https://arxiv.org/abs/2604.21375 — in: w2-stopping |
| S335 | https://arxiv.org/abs/2604.23783 — in: w2-stopping |
| S336 | https://arxiv.org/abs/2605.02801 — in: w2-stopping,ext-kimi-swarm-blog,w2-rltr,w2-orch-survey |
| S337 | https://arxiv.org/abs/2605.04984 — in: ext-credit-assignment |
| S338 | https://arxiv.org/abs/2605.10158 — in: ext-prm-vs-orm |
| S339 | https://arxiv.org/abs/2605.12484 — in: w2-matched-budget |
| S340 | https://arxiv.org/abs/2605.14483 — in: w2-stopping |
| S341 | https://arxiv.org/abs/2605.17292 — in: w2-stopping |
| S342 | https://arxiv.org/abs/2605.20873 — in: w2-planningbench,ext-planning-training |
| S343 | https://arxiv.org/abs/2605.22177 — in: w2-stopping |
| S344 | https://arxiv.org/abs/2605.27030 — in: ext-parallel-thinking |
| S345 | https://arxiv.org/abs/2606.00437 — in: ext-prm-vs-orm |
| S346 | https://arxiv.org/abs/2606.07027 — in: ext-prm-vs-orm |
| S347 | https://arxiv.org/abs/2606.08077 — in: ext-llm-judge-rubric |
| S348 | https://arxiv.org/abs/2606.09078 — in: ext-prm-vs-orm |
| S349 | https://arxiv.org/abs/2606.13040 — in: ext-prm-vs-orm |
| S350 | https://arxiv.org/abs/2606.13316 — in: w2-rltr |
| S351 | https://arxiv.org/abs/2606.22388 — in: ext-planning-training |
| S352 | https://arxiv.org/abs/2606.24525 — in: ext-prm-vs-orm |
| S353 | https://arxiv.org/abs/2606.26080 — in: ext-prm-vs-orm |
| S354 | https://arxiv.org/abs/2606.27009 — in: w2-stopping |
| S355 | https://arxiv.org/abs/2606.30613 — in: ext-codeact |
| S356 | https://arxiv.org/abs/2606.31484 — in: ext-parallel-thinking |
| S357 | https://arxiv.org/abs/2607.03991 — in: w2-stopping |
| S358 | https://arxiv.org/abs/2607.07435 — in: w2-rltr |
| S359 | https://arxiv.org/abs/2607.09153 — in: ext-prm-vs-orm |
| S360 | https://arxiv.org/abs/2607.11089 — in: w2-stopping |
| S361 | https://arxiv.org/abs/2607.13988 — in: ext-toolrl,ext-credit-assignment,w2-credit-detail |
| S362 | https://arxiv.org/abs/2607.14004 — in: w2-matched-budget |
| S363 | https://arxiv.org/abs/2607.24720 — in: ext-planning-training |
| S364 | https://arxiv.org/abs/2608.02009 — in: w2-stopping |
| S365 | https://arxiv.org/abs/2608.02276 — in: w2-matched-budget |
| S366 | https://arxiv.org/abs/2608.06113 — in: ext-planning-training |
| S367 | https://arxiv.org/abs/2608.06663 — in: ext-prm-vs-orm |
| S368 | https://arxiv.org/abs/2608.08020 — in: ext-parallel-thinking |
| S369 | https://arxiv.org/abs/2608.10178 — in: w2-matched-budget |
| S370 | https://arxiv.org/abs/2608.10357 — in: ext-toolrl |
| S371 | https://arxiv.org/abs/2608.13237 — in: w2-stopping |
| S372 | https://arxiv.org/abs/2608.16425 — in: ext-parallel-thinking |
| S373 | https://arxiv.org/abs/2608.18682 — in: ext-toolrl,w2-credit-detail |
| S374 | https://arxiv.org/abs/2608.18884 — in: w2-stopping |
| S375 | https://arxiv.org/abs/2608.22167 — in: ext-toolrl |
| S376 | https://arxiv.org/abs/2608.24588 — in: ext-credit-assignment |
| S377 | https://arxiv.org/abs/2608.27266 — in: w2-matched-budget |
| S378 | https://arxiv.org/html/2206.10498v4 — in: ext-benchmarks |
| S379 | https://arxiv.org/html/2305.20050 — in: ext-prm-vs-orm |
| S380 | https://arxiv.org/html/2312.08935 — in: ext-prm-vs-orm |
| S381 | https://arxiv.org/html/2402.16823 — in: ext-workflow-search |
| S382 | https://arxiv.org/html/2402.16823v3 — in: ext-workflow-search |
| S383 | https://arxiv.org/html/2402.19446 — in: ext-prm-vs-orm |
| S384 | https://arxiv.org/html/2402.19446v1 — in: ext-credit-assignment |
| S385 | https://arxiv.org/html/2406.04520v1 — in: ext-planning-training |
| S386 | https://arxiv.org/html/2406.06592 — in: ext-prm-vs-orm |
| S387 | https://arxiv.org/html/2406.11695 — in: ext-workflow-search |
| S388 | https://arxiv.org/html/2408.00764v3 — in: ext-curriculum-taskgen |
| S389 | https://arxiv.org/html/2408.07199 — in: ext-prm-vs-orm |
| S390 | https://arxiv.org/html/2408.08435 — in: ext-workflow-search |
| S391 | https://arxiv.org/html/2409.13373v1 — in: ext-planning-training |
| S392 | https://arxiv.org/html/2410.08146 — in: ext-prm-vs-orm |
| S393 | https://arxiv.org/html/2410.10762 — in: w2-matched-budget,ext-workflow-search |
| S394 | https://arxiv.org/html/2411.02337 — in: ext-webrl |
| S395 | https://arxiv.org/html/2411.02337v3 — in: ext-curriculum-taskgen |
| S396 | https://arxiv.org/html/2411.15124v2 — in: ext-rlvr-tulu |
| S397 | https://arxiv.org/html/2412.01981 — in: ext-prm-vs-orm |
| S398 | https://arxiv.org/html/2412.06559 — in: ext-prm-vs-orm |
| S399 | https://arxiv.org/html/2412.19437v2 — in: w2-slime-verl |
| S400 | https://arxiv.org/html/2412.21139 — in: ext-swe-rl |
| S401 | https://arxiv.org/html/2501.07301 — in: ext-prm-vs-orm |
| S402 | https://arxiv.org/html/2501.11425 — in: ext-credit-assignment |
| S403 | https://arxiv.org/html/2501.12599v2 — in: ext-kimi-k15 |
| S404 | https://arxiv.org/html/2501.12948 — in: ext-prm-vs-orm |
| S405 | https://arxiv.org/html/2502.01600 — in: ext-credit-assignment |
| S406 | https://arxiv.org/html/2502.10325 — in: ext-prm-vs-orm,ext-credit-assignment |
| S407 | https://arxiv.org/html/2502.16111v1 — in: ext-planning-training |
| S408 | https://arxiv.org/html/2502.18449v2 — in: w2-slime-verl |
| S409 | https://arxiv.org/html/2503.05592v2 — in: ext-search-rl |
| S410 | https://arxiv.org/html/2503.09516v3 — in: ext-search-rl |
| S411 | https://arxiv.org/html/2503.09572v1 — in: ext-planning-training |
| S412 | https://arxiv.org/html/2503.13657 — in: w2-mast |
| S413 | https://arxiv.org/html/2503.15478 — in: ext-prm-vs-orm,ext-credit-assignment |
| S414 | https://arxiv.org/html/2503.23383v1 — in: ext-toolrl |
| S415 | https://arxiv.org/html/2504.01400v3 — in: ext-toolrl |
| S416 | https://arxiv.org/html/2504.03160v2 — in: ext-search-rl |
| S417 | https://arxiv.org/html/2504.07164 — in: ext-swe-rl |
| S418 | https://arxiv.org/html/2504.11536v2 — in: ext-toolrl |
| S419 | https://arxiv.org/html/2504.12516v1 — in: ext-benchmarks |
| S420 | https://arxiv.org/html/2504.13837v2 — in: ext-rlvr-tulu |
| S421 | https://arxiv.org/html/2504.13958v1 — in: ext-toolrl |
| S422 | https://arxiv.org/html/2504.15257 — in: w2-flowreasoner-afm-verify |
| S423 | https://arxiv.org/html/2504.20073 — in: ext-prm-vs-orm |
| S424 | https://arxiv.org/html/2504.20073v2 — in: ext-credit-assignment |
| S425 | https://arxiv.org/html/2504.21798 — in: ext-swe-rl |
| S426 | https://arxiv.org/html/2504.21798v2 — in: ext-curriculum-taskgen |
| S427 | https://arxiv.org/html/2505.00024v2 — in: ext-toolrl |
| S428 | https://arxiv.org/html/2505.01441v1 — in: ext-toolrl |
| S429 | https://arxiv.org/html/2505.03335 — in: ext-curriculum-taskgen |
| S430 | https://arxiv.org/html/2505.07512v1 — in: ext-toolrl |
| S431 | https://arxiv.org/html/2505.10978 — in: ext-prm-vs-orm,ext-credit-assignment,w2-credit-detail |
| S432 | https://arxiv.org/html/2505.11821v2 — in: ext-credit-assignment |
| S433 | https://arxiv.org/html/2505.16410v1 — in: ext-toolrl |
| S434 | https://arxiv.org/html/2505.22312v2 — in: w2-slime-verl |
| S435 | https://arxiv.org/html/2505.22648v2 — in: ext-search-rl |
| S436 | https://arxiv.org/html/2505.23564 — in: ext-credit-assignment |
| S437 | https://arxiv.org/html/2506.03570 — in: ext-prm-vs-orm |
| S438 | https://arxiv.org/html/2506.10947v2 — in: ext-rlvr-tulu |
| S439 | https://arxiv.org/html/2506.13585v1 — in: w2-slime-verl |
| S440 | https://arxiv.org/html/2507.17307v4 — in: ext-credit-assignment |
| S441 | https://arxiv.org/html/2507.19457v2 — in: w2-matched-budget |
| S442 | https://arxiv.org/html/2507.19849 — in: ext-credit-assignment |
| S443 | https://arxiv.org/html/2507.20534 — in: ext-kimi-k2 |
| S444 | https://arxiv.org/html/2507.20534v1 — in: w2-slime-verl |
| S445 | https://arxiv.org/html/2507.20534v2 — in: ext-curriculum-taskgen |
| S446 | https://arxiv.org/html/2508.03680v1 — in: ext-frameworks |
| S447 | https://arxiv.org/html/2508.05004 — in: ext-curriculum-taskgen |
| S448 | https://arxiv.org/html/2508.06471v1 — in: w2-slime-verl |
| S449 | https://arxiv.org/html/2508.07976v1 — in: ext-search-rl |
| S450 | https://arxiv.org/html/2508.12685v3 — in: ext-toolrl |
| S451 | https://arxiv.org/html/2508.13167 — in: w2-flowreasoner-afm-verify |
| S452 | https://arxiv.org/html/2508.19598v1 — in: w2-rltr |
| S453 | https://arxiv.org/html/2509.02479v2 — in: ext-toolrl |
| S454 | https://arxiv.org/html/2509.02547 — in: ext-credit-assignment |
| S455 | https://arxiv.org/html/2509.02547v1 — in: ext-agentic-rl-survey |
| S456 | https://arxiv.org/html/2509.02547v5 — in: ext-agentic-rl-survey |
| S457 | https://arxiv.org/html/2509.19199v3 — in: ext-credit-assignment |
| S458 | https://arxiv.org/html/2509.21240 — in: ext-credit-assignment |
| S459 | https://arxiv.org/html/2509.25140v2 — in: ext-planning-training |
| S460 | https://arxiv.org/html/2510.08049 — in: ext-prm-vs-orm |
| S461 | https://arxiv.org/html/2511.08325 — in: ext-prm-vs-orm |
| S462 | https://arxiv.org/html/2511.10395v1 — in: ext-curriculum-taskgen |
| S463 | https://arxiv.org/html/2511.14846 — in: ext-credit-assignment |
| S464 | https://arxiv.org/html/2511.16108 — in: ext-swe-rl |
| S465 | https://arxiv.org/html/2512.17008 — in: ext-credit-assignment |
| S466 | https://arxiv.org/html/2512.18552v1 — in: w2-slime-verl |
| S467 | https://arxiv.org/html/2602.02276 — in: ext-kimi-k25 |
| S468 | https://arxiv.org/html/2602.02276v1 — in: w2-openparl,ext-curriculum-taskgen |
| S469 | https://arxiv.org/html/2602.02276v2 — in: ext-kimi-swarm-blog,w2-parl-paper |
| S470 | https://arxiv.org/html/2602.02276v2/pa-rl-progress.png — in: w2-parl-paper |
| S471 | https://arxiv.org/html/2602.03845v2 — in: w2-util-metrics |
| S472 | https://arxiv.org/html/2602.04634v1 — in: w2-openparl |
| S473 | https://arxiv.org/html/2602.11767 — in: ext-credit-assignment |
| S474 | https://arxiv.org/html/2603.06194v1 — in: ext-credit-assignment |
| S475 | https://arxiv.org/html/2603.08754v1 — in: ext-credit-assignment |
| S476 | https://arxiv.org/html/2603.21972 — in: ext-prm-vs-orm |
| S477 | https://arxiv.org/html/2604.09459 — in: ext-credit-assignment |
| S478 | https://arxiv.org/html/2604.18292v1 — in: ext-curriculum-taskgen |
| S479 | https://arxiv.org/html/2605.02801v1 — in: w2-stopping,w2-orch-survey,ext-credit-assignment |
| S480 | https://arxiv.org/html/2605.04984v1 — in: ext-credit-assignment |
| S481 | https://arxiv.org/html/2605.10158 — in: ext-prm-vs-orm |
| S482 | https://arxiv.org/html/2605.20873 — in: w2-planningbench |
| S483 | https://arxiv.org/html/2605.20873v2 — in: ext-planning-training |
| S484 | https://arxiv.org/html/2606.00437 — in: ext-prm-vs-orm |
| S485 | https://arxiv.org/html/2606.07027 — in: ext-prm-vs-orm |
| S486 | https://arxiv.org/html/2606.09078 — in: ext-prm-vs-orm |
| S487 | https://arxiv.org/html/2606.13040 — in: ext-prm-vs-orm |
| S488 | https://arxiv.org/html/2606.24525 — in: ext-prm-vs-orm |
| S489 | https://arxiv.org/html/2606.26080 — in: ext-prm-vs-orm |
| S490 | https://arxiv.org/html/2607.07435v1 — in: w2-rltr |
| S491 | https://arxiv.org/html/2607.09153 — in: ext-prm-vs-orm |
| S492 | https://arxiv.org/html/2607.13988 — in: ext-credit-assignment,w2-credit-detail |
| S493 | https://arxiv.org/html/2607.14004 — in: w2-matched-budget |
| S494 | https://arxiv.org/html/2608.06663 — in: ext-prm-vs-orm |
| S495 | https://arxiv.org/html/2608.18682 — in: w2-credit-detail |
| S496 | https://arxiv.org/html/2608.24588v2 — in: ext-credit-assignment |
| S497 | https://arxiv.org/html/2608.25683v1 — in: ext-credit-assignment |
| S498 | https://arxiv.org/html/2608.27266 — in: w2-matched-budget |
| S499 | https://arxiv.org/pdf/2411.02337 — in: ext-webrl |
| S500 | https://arxiv.org/pdf/2502.18449v1 — in: ext-swe-rl |
| S501 | https://arxiv.org/pdf/2503.13657 — in: w2-mast |
| S502 | https://arxiv.org/pdf/2503.13657v1 — in: w2-mast |
| S503 | https://arxiv.org/pdf/2602.02276 — in: w2-parl-paper |
| S504 | https://arxiv.org/pdf/2605.02801 — in: w2-orch-survey |
| S505 | https://arxiv.org/pdf/2605.20873 — in: w2-planningbench |
| S506 | https://arxiv.org/search/ — in: w2-crosstask,w2-orch-survey |
| S507 | https://arxiv.org/search/?query=%22curriculum%22+%22reinforcement+learning%22+LLM+reasoning&searchtype=all — in: ext-kimi-k15 |
| S508 | https://arxiv.org/search/?query=%22Dr.+GRPO%22+OR+%22biased+GRPO%22+length+bias&searchtype=all — in: ext-kimi-k15 |
| S509 | https://arxiv.org/search/?query=%22length+penalty%22+reasoning+RL&searchtype=all — in: ext-kimi-k15 |
| S510 | https://arxiv.org/search/?query=%22long-CoT%22+reinforcement+learning&searchtype=all — in: ext-kimi-k15 |
| S511 | https://arxiv.org/search/?query=%22online+mirror+descent%22+language+model&searchtype=all — in: ext-kimi-k15 |
| S512 | https://arxiv.org/search/?query=%22tool-use+completeness%22&searchtype=all — in: w2-rltr |
| S513 | https://arxiv.org/search/?query=ASearcher&searchtype=all — in: ext-search-rl |
| S514 | https://arxiv.org/search/?query=DAPO+decoupled+clip+AND+dynamic+sampling&searchtype=all — in: ext-kimi-k15 |
| S515 | https://arxiviq.substack.com/p/gepa-reflective-prompt-evolution — in: w2-matched-budget |
| S516 | https://australiansciencejournals.com/ajdsa/article/download/3667/4583 — in: w2-crosstask |
| S517 | https://australiansciencejournals.com/bigdata/article/download/3639/4558 — in: w2-crosstask |
| S518 | https://blog.langchain.com/context-engineering/ — in: ext-cognition-counter |
| S519 | https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf — in: ext-openai-agents |
| S520 | https://cognition.ai/blog/dont-build-multi-agents — in: ext-cognition-counter |
| S521 | https://datasets-server.huggingface.co/rows?dataset=gaia-benchmark%2Fresults_public&config=2023&split=test&offset=0&length=100 — in: ext-benchmarks |
| S522 | https://dblp.org/search/publ/api?q=PlanningBench — in: w2-planningbench |
| S523 | https://dblp.org/search/publ/api?q=reward%20modeling%20rubrics — in: ext-llm-judge-rubric |
| S524 | https://developers.openai.com/api/docs/changelog — in: ext-openai-agents |
| S525 | https://docs.primeintellect.ai/hosted-training/environment-model — in: ext-env-infra |
| S526 | https://docs.primeintellect.ai/llms.txt — in: ext-env-infra |
| S527 | https://docs.primeintellect.ai/sandboxes/overview — in: ext-env-infra |
| S528 | https://docs.primeintellect.ai/tutorials-environments/create — in: ext-env-infra |
| S529 | https://docs.primeintellect.ai/tutorials-environments/environments — in: ext-env-infra |
| S530 | https://docs.primeintellect.ai/tutorials-environments/getting-started — in: ext-env-infra |
| S531 | https://docs.primeintellect.ai/verifiers/v1/env — in: ext-env-infra |
| S532 | https://docs.primeintellect.ai/verifiers/v1/harbor — in: ext-env-infra |
| S533 | https://docs.skyrl.ai/docs/getting-started/inference_architecture — in: ext-frameworks |
| S534 | https://docs.skyrl.ai/docs/getting-started/overview — in: ext-frameworks |
| S535 | https://docs.skyrl.ai/docs/recipes/overview — in: ext-swe-rl |
| S536 | https://docs.skyrl.ai/docs/tutorials/fully_async — in: ext-frameworks |
| S537 | https://docs.skyrl.ai/docs/tutorials/one_step_off_async — in: ext-frameworks |
| S538 | https://doi.org/10.1145/3786335.3813164 — in: w2-matched-budget |
| S539 | https://doi.org/10.1145/3786335.3813167 — in: w2-matched-budget |
| S540 | https://doi.org/10.71465/ajbd3639 — in: w2-crosstask |
| S541 | https://doi.org/10.71465/ajdsa3667 — in: w2-crosstask |
| S542 | https://dspy.ai/api/optimizers/GEPA/overview/ — in: w2-matched-budget |
| S543 | https://dspy.ai/getting-started/gepa-optimization/ — in: ext-workflow-search |
| S544 | https://e2b.dev/blog/up-to-5x-faster-sandboxes — in: ext-env-infra |
| S545 | https://e2b.dev/llms.txt — in: ext-env-infra |
| S546 | https://e2b.dev/pricing — in: ext-env-infra |
| S547 | https://en.wikipedia.org/wiki/ChatGPT_Deep_Research — in: ext-openai-agents |
| S548 | https://en.wikipedia.org/wiki/DeepSeek_R1 — in: w2-slime-verl |
| S549 | https://en.wikipedia.org/wiki/Kimi_(AI — in: ext-kimi-researcher |
| S550 | https://en.wikipedia.org/wiki/Moonshot_AI — in: w2-kimi-k26-k3 |
| S551 | https://entropytown.com/articles/2025-11-07-kimi-k2-thinking/ — in: ext-kimi-researcher |
| S552 | https://export.arxiv.org/api/query — in: w2-crosstask,w2-credit-detail |
| S553 | https://forgecode.dev/blog/kimi-k2-vs-sonnet-4-vs-gemini-2.5-pro/ — in: ext-kimi-community |
| S554 | https://gepa-ai.github.io/gepa/ — in: w2-matched-budget |
| S555 | https://gepa-ai.github.io/gepa/api/adapters/TerminalBenchAdapter/ — in: w2-matched-budget |
| S556 | https://gepa-ai.github.io/gepa/blog/2026/02/18/automatically-learning-skills-for-coding-agents/ — in: w2-matched-budget |
| S557 | https://gepa-ai.github.io/gepa/blog/2026/02/18/introducing-optimize-anything/ — in: w2-matched-budget |
| S558 | https://gist.github.com/aarvay/3c930cb4d3c6696409af2b31d4953294 — in: ext-kimi-community |
| S559 | https://github.com/AMAP-ML/Tree-GRPO — in: ext-parallel-thinking |
| S560 | https://github.com/anthropics/anthropic-cookbook — in: ext-anthropic-multiagent |
| S561 | https://github.com/apple/ml-loop — in: ext-credit-assignment |
| S562 | https://github.com/ByteDance-Seed/WideSearch — in: w2-openparl |
| S563 | https://github.com/chanwoo-park-official/MAPoRL — in: ext-multiagent-rl |
| S564 | https://github.com/Dao-AILab/gambit-parallel-reasoning — in: ext-parallel-thinking |
| S565 | https://github.com/data-for-agents/insta — in: ext-webrl |
| S566 | https://github.com/digirl-agent/digirl — in: ext-webrl |
| S567 | https://github.com/facebookresearch/sweet_rl — in: ext-credit-assignment |
| S568 | https://github.com/facebookresearch/threadweaver — in: ext-parallel-thinking |
| S569 | https://github.com/FoundationAgents/AFlow — in: ext-multiagent-rl,ext-workflow-search |
| S570 | https://github.com/Gen-Verse/ScoreFlow — in: ext-multiagent-rl |
| S571 | https://github.com/gepa-ai/gepa — in: ext-workflow-search |
| S572 | https://github.com/GuanxingLu/OpenPARL — in: ext-kimi-swarm-blog |
| S573 | https://github.com/huggingface/OpenEnv — in: ext-env-infra |
| S574 | https://github.com/langfengQ/verl-agent — in: ext-credit-assignment |
| S575 | https://github.com/MASWorks/MAS-GPT — in: ext-multiagent-rl |
| S576 | https://github.com/metauto-ai/GPTSwarm — in: ext-workflow-search |
| S577 | https://github.com/MoonshotAI/checkpoint-engine — in: ext-kimi-k2 |
| S578 | https://github.com/MoonshotAI/K2-Vendor-Verfier — in: ext-kimi-community |
| S579 | https://github.com/MoonshotAI/Kimi-K2 — in: ext-kimi-researcher,ext-kimi-k2 |
| S580 | https://github.com/MoonshotAI/Kimi-K2.5 — in: ext-kimi-k25,w2-parl-paper,ext-kimi-community |
| S581 | https://github.com/MoonshotAI/Kimi-Researcher — in: ext-kimi-researcher |
| S582 | https://github.com/openai/openai-agents-python — in: ext-openai-agents |
| S583 | https://github.com/OpenBMB/ChatDev/tree/puppeteer — in: ext-multiagent-rl |
| S584 | https://github.com/Parallel-Reasoning/APR — in: ext-parallel-thinking |
| S585 | https://github.com/qiancheng0/ToolRL — in: ext-toolrl |
| S586 | https://github.com/RLinf/RLinf — in: w2-openparl |
| S587 | https://github.com/Rulin3/Spurious-Reward — in: ext-rlvr-tulu |
| S588 | https://github.com/SafwanAlselwi/LLM-RL — in: ext-agentic-rl-survey |
| S589 | https://github.com/sail-sg/FlowReasoner — in: ext-multiagent-rl |
| S590 | https://github.com/sanjibanc/agent_prm — in: ext-credit-assignment |
| S591 | https://github.com/ShengranHu/ADAS — in: ext-workflow-search |
| S592 | https://github.com/sierra-research/tau2-bench.git — in: ext-env-infra |
| S593 | https://github.com/stanfordnlp/dspy — in: ext-workflow-search |
| S594 | https://github.com/Tencent-Hunyuan/PlanningBench — in: w2-planningbench |
| S595 | https://github.com/THUDM/slime — in: w2-slime-verl,ext-glm-qwen |
| S596 | https://github.com/THUDM/VisualAgentBench — in: ext-webrl |
| S597 | https://github.com/THUDM/VisualAgentBench/tree/main/VAB-WebArena-Lite — in: ext-webrl |
| S598 | https://github.com/THUDM/WebRL — in: ext-webrl |
| S599 | https://github.com/thunlp/Optima — in: ext-multiagent-rl |
| S600 | https://github.com/ventr1c/Awesome-RL-based-Agentic-Search-Papers — in: ext-agentic-rl-survey |
| S601 | https://github.com/verl-project/uni-agent — in: w2-slime-verl |
| S602 | https://github.com/volcengine/verl — in: ext-frameworks,w2-slime-verl |
| S603 | https://github.com/web-arena-x/webarena — in: ext-webrl |
| S604 | https://github.com/xxzcc/Awesome-Credit-Assignment-in-LLM-RL — in: ext-credit-assignment |
| S605 | https://github.com/xxzcc/awesome-llm-mas-rl — in: w2-orch-survey |
| S606 | https://github.com/yanweiyue/masrouter — in: ext-multiagent-rl |
| S607 | https://github.com/zhengkid/Parallel-R1 — in: ext-parallel-thinking |
| S608 | https://huggingface.co/api/models?author=moonshotai — in: ext-kimi-k25 |
| S609 | https://huggingface.co/api/models/moonshotai/Kimi-K2-Thinking — in: ext-kimi-k25 |
| S610 | https://huggingface.co/api/models/moonshotai/Kimi-K2.5 — in: ext-kimi-k25 |
| S611 | https://huggingface.co/api/papers/2309.12499 — in: ext-codeact |
| S612 | https://huggingface.co/api/papers/2508.19598 — in: w2-rltr |
| S613 | https://huggingface.co/api/papers/search — in: ext-planning-training |
| S614 | https://huggingface.co/api/papers/search?q=optimal%20tool%20calls%20reinforcement — in: w2-rltr |
| S615 | https://huggingface.co/api/papers/search?q=plan-as-code — in: ext-codeact |
| S616 | https://huggingface.co/api/papers/search?q=plan+code+agent — in: ext-codeact |
| S617 | https://huggingface.co/api/papers/search?q=rubrics+as+rewards — in: ext-llm-judge-rubric |
| S618 | https://huggingface.co/api/spaces?author=openenv — in: ext-env-infra |
| S619 | https://huggingface.co/datasets/ByteDance-Seed/WideSearch — in: w2-openparl |
| S620 | https://huggingface.co/datasets/gaia-benchmark/GAIA — in: ext-benchmarks |
| S621 | https://huggingface.co/datasets/google-research-datasets/natural_plan — in: ext-benchmarks |
| S622 | https://huggingface.co/datasets/inclusionAI/ASearcher-Local-Knowledge — in: w2-openparl |
| S623 | https://huggingface.co/datasets/mcemri/MAST-Data — in: w2-mast |
| S624 | https://huggingface.co/datasets/openai/browsecomp — in: ext-benchmarks |
| S625 | https://huggingface.co/datasets/RLinf/WideSeek-R1-train-data — in: w2-openparl |
| S626 | https://huggingface.co/datasets/tencent/PlanningBench — in: w2-planningbench |
| S627 | https://huggingface.co/datasets/tencent/PlanningBench/resolve/main/data/PlanningBench-eval.jsonl — in: w2-planningbench |
| S628 | https://huggingface.co/datasets/WideSeek-R1/Wiki-2018-Corpus — in: w2-openparl |
| S629 | https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 — in: ext-glm-qwen |
| S630 | https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp — in: ext-glm-qwen |
| S631 | https://huggingface.co/docs/openenv/environments — in: ext-env-infra |
| S632 | https://huggingface.co/MiniMaxAI/MiniMax-M2 — in: ext-glm-qwen |
| S633 | https://huggingface.co/moonshotai/Kimi-K2-Instruct — in: ext-kimi-k2 |
| S634 | https://huggingface.co/moonshotai/Kimi-K2-Thinking — in: ext-kimi-researcher,ext-kimi-community |
| S635 | https://huggingface.co/moonshotai/Kimi-K2-Thinking/raw/main/README.md — in: w2-kimi-k26-k3 |
| S636 | https://huggingface.co/moonshotai/Kimi-K2.5 — in: ext-kimi-k25 |
| S637 | https://huggingface.co/moonshotai/Kimi-K2.6/raw/main/README.md — in: w2-kimi-k26-k3 |
| S638 | https://huggingface.co/moonshotai/Kimi-K3/raw/main/README.md — in: w2-kimi-k26-k3 |
| S639 | https://huggingface.co/papers?q=RLVR+generalization+transfer — in: ext-transfer-evidence |
| S640 | https://huggingface.co/papers/2605.20873 — in: w2-planningbench |
| S641 | https://huggingface.co/papers/2607.13988 — in: w2-credit-detail |
| S642 | https://huggingface.co/papers/2608.18682 — in: w2-credit-detail |
| S643 | https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct — in: ext-glm-qwen |
| S644 | https://huggingface.co/spaces/gaia-benchmark/leaderboard/raw/main/app.py — in: ext-benchmarks |
| S645 | https://huggingface.co/spaces/muset-ai/DeepResearch-Bench-Leaderboard/raw/main/data/leaderboard.csv — in: ext-benchmarks |
| S646 | https://huggingface.co/THUDM/webrl-glm-4-9b — in: ext-webrl |
| S647 | https://huggingface.co/zai-org/GLM-4.5 — in: ext-glm-qwen |
| S648 | https://iclr.cc/virtual/2026/poster/10008772 — in: ext-credit-assignment |
| S649 | https://lmsys.org/blog/2025-07-09-slime/ — in: ext-glm-qwen |
| S650 | https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus — in: ext-cognition-counter,w2-context-bounds |
| S651 | https://manus.im/blog/manus-1.5-release — in: w2-context-bounds |
| S652 | https://manus.im/blog/manus-wide-research-solve-context-problem — in: w2-context-bounds |
| S653 | https://modal.com/docs/examples/harbor_evals.md — in: ext-env-infra |
| S654 | https://modal.com/docs/guide/sandboxes.md — in: ext-env-infra |
| S655 | https://modal.com/llms.txt — in: ext-env-infra |
| S656 | https://moonshotai.github.io/Kimi-K2/ — in: ext-kimi-k2 |
| S657 | https://moonshotai.github.io/Kimi-K2/thinking.html — in: ext-kimi-researcher,w2-kimi-k26-k3 |
| S658 | https://moonshotai.github.io/Kimi-Researcher/ — in: ext-kimi-researcher,ext-credit-assignment |
| S659 | https://namu.wiki/w/Kimi — in: w2-kimi-k26-k3 |
| S660 | https://news.ycombinator.com/item?id=44535078 — in: ext-frameworks,ext-llm-judge-rubric |
| S661 | https://news.ycombinator.com/item?id=45096962 — in: ext-cognition-counter |
| S662 | https://news.ycombinator.com/item?id=45836070 — in: ext-kimi-researcher,ext-kimi-community |
| S663 | https://news.ycombinator.com/item?id=46775961 — in: ext-kimi-community |
| S664 | https://news.ycombinator.com/item?id=46826597 — in: ext-kimi-community |
| S665 | https://news.ycombinator.com/item?id=47126614 — in: ext-kimi-community |
| S666 | https://news.ycombinator.com/item?id=47452404 — in: ext-kimi-community |
| S667 | https://news.ycombinator.com/item?id=47835735 — in: ext-kimi-community |
| S668 | https://news.ycombinator.com/item?id=49007610 — in: ext-kimi-community |
| S669 | https://observablehq.com/@tomlarkworthy/gepa — in: w2-matched-budget |
| S670 | https://old.reddit.com/r/singularity/comments/1ryrs2w/ — in: ext-kimi-community |
| S671 | https://openai.com/index/browsecomp/ — in: ext-benchmarks |
| S672 | https://openai.com/index/introducing-agentkit/ — in: ext-openai-agents |
| S673 | https://openai.com/index/introducing-deep-research/ — in: ext-benchmarks |
| S674 | https://openai.com/news/rss.xml — in: ext-openai-agents |
| S675 | https://openai.com/sitemap.xml — in: ext-openai-agents |
| S676 | https://openpipe.ai/blog/reward-hacking — in: ext-llm-judge-rubric |
| S677 | https://openpipe.ai/blog/ruler — in: ext-llm-judge-rubric |
| S678 | https://openreview.net/forum?id=a7Qa4CcHak — in: ext-benchmarks |
| S679 | https://openreview.net/forum?id=ooROvpmxMV — in: ext-credit-assignment |
| S680 | https://qwenlm.github.io/blog/qwen3-coder/ — in: ext-glm-qwen |
| S681 | https://raw.githubusercontent.com/gepa-ai/gepa/main/README.md — in: w2-matched-budget |
| S682 | https://raw.githubusercontent.com/GuanxingLu/OpenPARL/main/BLOG.md — in: w2-openparl |
| S683 | https://raw.githubusercontent.com/MoonshotAI/Kimi-K2.5/master/README.md — in: w2-parl-paper,ext-curriculum-taskgen |
| S684 | https://research.ibm.com/publications/tsr-trajectorysearch-rollouts-for-multiturn-rl-of-llm-agents — in: ext-credit-assignment |
| S685 | https://research.trychroma.com/context-rot — in: ext-cognition-counter,w2-context-bounds |
| S686 | https://rlancemartin.github.io/2025/10/15/manus/ — in: ext-cognition-counter |
| S687 | https://rlinf.readthedocs.io/en/latest/rst_source/examples/agentic/wideseek_r1/index.html — in: w2-openparl |
| S688 | https://routerlab.ch/blog/kimi-k2-5 — in: ext-kimi-swarm-blog |
| S689 | https://sierra.ai/blog/benchmarking-agents-in-collaborative-real-world-scenarios — in: ext-benchmarks |
| S690 | https://simonwillison.net/2025/Jun/27/context-engineering/ — in: ext-cognition-counter |
| S691 | https://sites.google.com/berkeley.edu/mast/ — in: w2-mast |
| S692 | https://taubench.com/leaderboard?benchmark=core — in: ext-benchmarks |
| S693 | https://techcrunch.com/2025/01/23/openai-launches-operator-an-ai-agent-that-performs-tasks-autonomously/ — in: ext-openai-agents |
| S694 | https://techcrunch.com/2025/02/02/openai-unveils-a-new-chatgpt-agent-for-deep-research/ — in: ext-openai-agents |
| S695 | https://techcrunch.com/2025/04/16/openai-launches-a-pair-of-ai-reasoning-models-o3-and-o4-mini/ — in: ext-openai-agents |
| S696 | https://techcrunch.com/2025/05/16/openai-launches-codex-an-ai-coding-agent-in-chatgpt/ — in: ext-openai-agents |
| S697 | https://techcrunch.com/2025/07/17/openai-launches-a-general-purpose-agent-in-chatgpt/ — in: ext-openai-agents |
| S698 | https://techcrunch.com/2025/08/07/openais-gpt-5-is-here/ — in: ext-openai-agents |
| S699 | https://techcrunch.com/2025/10/06/openai-launches-agentkit-to-help-developers-build-and-ship-ai-agents/ — in: ext-openai-agents |
| S700 | https://techcrunch.com/2025/10/06/openai-launches-apps-inside-of-chatgpt/ — in: ext-openai-agents |
| S701 | https://techcrunch.com/2025/10/06/openai-ramps-up-developer-push-with-more-powerful-models-in-its-api/ — in: ext-openai-agents |
| S702 | https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/ — in: w2-kimi-k26-k3 |
| S703 | https://the-decoder.com/moonshot-ai-releases-kimi-k2-5-claims-most-powerful-open-weight-model-with-100-agent-coordination/ — in: w2-kimi-k26-k3 |
| S704 | https://the-decoder.com/moonshot-ai-releases-kimi-k3-open-weights-and-infrastructure-after-shaking-up-the-frontier-model-race/ — in: w2-kimi-k26-k3 |
| S705 | https://the-decoder.com/moonshot-pauses-new-kimi-k3-subscriptions-after-gpu-demand-maxes-out-in-48-hours/ — in: w2-kimi-k26-k3 |
| S706 | https://the-decoder.com/open-weight-kimi-k2-6-takes-on-gpt-5-4-and-claude-opus-4-6-with-agent-swarms/ — in: w2-kimi-k26-k3 |
| S707 | https://thezvi.substack.com/p/kimi-k2 — in: ext-kimi-community |
| S708 | https://thezvi.substack.com/p/kimi-k25 — in: ext-kimi-community |
| S709 | https://twitter.com/fynnso/status/2034706304875602030 — in: ext-kimi-community |
| S710 | https://twitter.com/leerob/status/2035050444347600936 — in: ext-kimi-community |
| S711 | https://verl.readthedocs.io/en/latest/sglang_multiturn/sandbox_fusion.html — in: ext-env-infra |
| S712 | https://web.archive.org/web/20260614220820/https://cognition.ai/blog/dont-build-multi-agents — in: ext-cognition-counter |
| S713 | https://web.archive.org/web/20260718095914/https://novasky-ai.notion.site/skyrl-v0 — in: ext-swe-rl |
| S714 | https://www.alphaxiv.org/abs/2505.10978 — in: w2-credit-detail |
| S715 | https://www.alphaxiv.org/abs/2605.20873 — in: w2-planningbench |
| S716 | https://www.anthropic.com/engineering/building-c-compiler — in: w2-orch-survey |
| S717 | https://www.anthropic.com/engineering/built-multi-agent-research-system — in: ext-cognition-counter,w2-util-metrics |
| S718 | https://www.anthropic.com/engineering/claude-code-best-practices — in: ext-cognition-counter,ext-anthropic-multiagent |
| S719 | https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents — in: ext-anthropic-multiagent |
| S720 | https://www.anthropic.com/engineering/eval-awareness-browsecomp — in: ext-anthropic-multiagent |
| S721 | https://www.anthropic.com/engineering/multi-agent-research-system — in: ext-benchmarks,w2-redteam,ext-anthropic-multiagent |
| S722 | https://www.anthropic.com/features/project-deal — in: w2-redteam |
| S723 | https://www.anthropic.com/research/building-effective-agents — in: ext-anthropic-multiagent |
| S724 | https://www.anthropic.com/research/glasswing-initial-update — in: w2-redteam |
| S725 | https://www.anthropic.com/research/multiagent-systems — in: w2-redteam,ext-cognition-counter |
| S726 | https://www.anthropic.com/research/team/frontier-red-team — in: w2-redteam |
| S727 | https://www.anthropic.com/sitemap.xml — in: w2-redteam,ext-anthropic-multiagent |
| S728 | https://www.bing.com/search — in: ext-kimi-community |
| S729 | https://www.databricks.com/blog/building-state-art-enterprise-agents-90x-cheaper-automated-prompt-optimization — in: w2-matched-budget |
| S730 | https://www.daytona.io/docs/en/isolation.md — in: ext-env-infra |
| S731 | https://www.daytona.io/docs/en/limits.md — in: ext-env-infra |
| S732 | https://www.daytona.io/docs/en/sandboxes.md — in: ext-env-infra |
| S733 | https://www.daytona.io/docs/en/snapshots.md — in: ext-env-infra |
| S734 | https://www.daytona.io/docs/en/warm-pools.md — in: ext-env-infra |
| S735 | https://www.interconnects.ai/p/kimi-k2-and-when-deepseek-moments — in: ext-kimi-community |
| S736 | https://www.interconnects.ai/p/kimi-k2-thinking-what-it-means — in: ext-kimi-researcher |
| S737 | https://www.kimi.ai/ai-models/kimi-k2-6 — in: w2-kimi-k26-k3 |
| S738 | https://www.kimi.ai/ai-models/kimi-k3 — in: w2-kimi-k26-k3 |
| S739 | https://www.kimi.ai/blog/agent-swarm — in: ext-kimi-swarm-blog,int-thread-context |
| S740 | https://www.kimi.com/blog/ — in: w2-kimi-k26-k3 |
| S741 | https://www.kimi.com/blog/kimi-k2-5.html — in: ext-kimi-k25,w2-parl-paper,w2-orch-survey,ext-kimi-community |
| S742 | https://www.kimi.com/blog/kimi-k2-6 — in: ext-kimi-community |
| S743 | https://www.kimi.com/blog/kimi-k2-6.html — in: ext-kimi-k25,w2-kimi-k26-k3 |
| S744 | https://www.kimi.com/blog/kimi-k2-thinking.html — in: w2-kimi-k26-k3 |
| S745 | https://www.kimi.com/blog/kimi-k3 — in: w2-kimi-k26-k3 |
| S746 | https://www.kimi.com/en — in: w2-kimi-k26-k3 |
| S747 | https://www.kimi.com/en/agent-swarm — in: ext-kimi-swarm-blog |
| S748 | https://www.letta.com/blog/benchmarking-ai-agent-memory/ — in: w2-context-bounds |
| S749 | https://www.letta.com/blog/guide-to-context-engineering/ — in: w2-context-bounds |
| S750 | https://www.letta.com/blog/letta-leaderboard/ — in: w2-context-bounds |
| S751 | https://www.letta.com/blog/letta-v1-agent/ — in: w2-context-bounds |
| S752 | https://www.letta.com/blog/sleep-time-compute — in: ext-cognition-counter,w2-context-bounds |
| S753 | https://www.minimax.io/models/text/m25 — in: ext-glm-qwen |
| S754 | https://www.minimax.io/news/minimax-m2 — in: ext-glm-qwen |
| S755 | https://www.minimax.io/news/minimaxm1 — in: w2-slime-verl |
| S756 | https://www.mojeek.com/search — in: ext-kimi-community |
| S757 | https://www.reddit.com/r/LocalLLaMA/ — in: ext-kimi-community |
| S758 | https://www.swebench.com/ — in: ext-benchmarks |
| S759 | https://www.swebench.com/verified.html — in: ext-benchmarks |
| S760 | https://www.tbench.ai/benchmarks — in: ext-benchmarks |