RLVR 답변이 알고리즘에 따라 길어지거나 무너지는 이유: LUSPO
LUSPO가 RLVR objective의 sequence-length gradient 편향을 normalization해 GRPO의 장문화와 GSPO의 collapse를 줄이는 원리, length bucket, reasoning quality, token 비용 검증법을 설명합니다.
LUSPO가 RLVR objective의 sequence-length gradient 편향을 normalization해 GRPO의 장문화와 GSPO의 collapse를 줄이는 원리, length bucket, reasoning quality, token 비용 검증법을 설명합니다.
Context Forcing이 long-context teacher와 sink, slow, fast KV memory로 긴 rollout student의 supervision mismatch를 줄이는 원리, ECL 20초와 1분 생성의 차이, 비용, drift 검증법을 설명합니다.
MPRM data가 포화되는 이유와 BIS가 label mixture, Monte Carlo reliability로 정보량 높은 10%를 고르는 원리, annotation 비용, rare pattern, distribution shift 검증법을 설명합니다.
OmniSIFT가 STVP로 video 중복을 줄이고 visual anchor로 audio token을 고르는 비대칭 압축 원리, 75% 제거 조건과 audio-first 사건, STE, edge 비용 검증법을 설명합니다.
3DiMo가 body, hand motion token, perspective augmentation과 annealed SMPL supervision으로 cross-view human video를 제어하는 원리, geometry 오류, 극단 시점, 평가 조건을 설명합니다.
Green-VLA의 L0~R2 staged training, flow-matching action expert와 OOD, episode guidance를 설명하고 15~20% 주장, RL 비용, latency, 안전 검증 조건을 점검합니다.
CodeOCR이 syntax-highlighted code image로 최대 8배 token 압축을 시험한 원리, clone detection과 completion의 차이, symbol 오독, render 비용과 hybrid routing 기준을 설명합니다.
PaperBanana가 reference retrieval, content planning, neural/code rendering, VLM critique로 학술 figure를 만드는 과정, 35% 결과의 범위와 수치, 출처, 검수 조건을 설명합니다.
LingBot-VA가 video, action token을 causal sequence로 학습하고 미래 visual state에서 action을 복원하는 원리, counterfactual 검증, prediction drift, latency와 안전 경계를 설명합니다.
VTC-R1이 이전 reasoning을 image optical memory로 바꿔 3.4배 token 압축, 2.7배 latency 개선을 보고한 원리, segment 설계, OCR 오류, text fallback 기준을 설명합니다.