From c497e3a32f4db5a5b8539277329f4ca47c426bac Mon Sep 17 00:00:00 2001 From: Quang Bui Date: Wed, 29 Jul 2026 07:09:18 +0700 Subject: [PATCH] =?UTF-8?q?paper:=20fold=20in=20the=20team's=20AAAI=5FFina?= =?UTF-8?q?l=20pass=20=E2=80=94=20figures=20wired,=20camera-ready=20author?= =?UTF-8?q?s=20staged?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The team's AAAI_Final/ folder was the pre-audit paper plus two intents: an author block to fill in, and the split fig_repair panels restored. Both are honored on top of the corrected text (so none of the 42 number fixes are lost): the accuracy panel now sits in the main-text repair section (fig:repair-acc, 0.8 columnwidth) with the K-scaling panel replacing the combined figure in the appendix, and the named equal-contribution author block is staged as a comment above the active 'Anonymous submission' line — tomorrow's submission is double-blind, so it swaps in only at camera-ready. The team's emptied \title{} was clearly accidental and the title is kept. Fitting the new main-text figure back into the 7-page budget took a prose pass: one genuinely duplicated clause removed (the linear rows were called mechanism-evidence twice), the geometry paragraph no longer restates its own premise, and the conclusion tail tightened. Content ends exactly at the bottom of page 7, references start page 8, no dangling refs, verifier 35/35. AAAI_Final/ deleted after the merge as requested. --- paper/figures/fig_repair_accuracy.pdf | Bin 0 -> 14516 bytes paper/figures/fig_repair_kscaling.pdf | Bin 0 -> 14288 bytes paper/tex/marc_aaai.tex | 85 ++++++++++++++------------ paper/tex/marc_aaai_appendix.tex | 14 ++--- 4 files changed, 54 insertions(+), 45 deletions(-) create mode 100644 paper/figures/fig_repair_accuracy.pdf create mode 100644 paper/figures/fig_repair_kscaling.pdf diff --git a/paper/figures/fig_repair_accuracy.pdf b/paper/figures/fig_repair_accuracy.pdf new file mode 100644 index 0000000000000000000000000000000000000000..c681293ab27748134b5d884a33901c1e1f40fff1 GIT binary patch literal 14516 zcmb_@2{=_<)PF+hk_=G_9fXW`zFbr0GIOs?QKoAsQ}qAUTf`j*81&z)|E8Z)mK0(;t-OhgHTBg z0u7-cU*|mtH8lundBBeXA$6SSPF}w55YpVqokD{!fPe&PXdoy)t}vnUFCFxKedrLD zL4dR|bKXO7p+iaww|ar}ZDhI=odV$&8qA&ObPCl6!o#-+B-z8s)zilvA}l_o`nr%Q zbZ94-R#zXeLJ6cpNCR)cg7)H7XYpzPIsC#7`gZ|fzG3dE6u{jA`$!Un=Ic*&0rJ57 zgZWJ;uAWZXzJXvy6!^zt(TW%dgTpB*gW-`nKrqmo1`!w7)AI501$Ste^jBrTkAG;- znBwD3_khrgGwOSK0eK;$z86q}F2%*yl>)O(qf;qP-iW}gYt|-Y2_>FuW7eY{+^HWF z?uf1AShZtw%G!o2#PL_CybbjJf!d;NowRl zUvue~vFEXK!6s7sa69?6rlzZNvmQTtXUu-Bek8&#b~2h2YM`dD-;A{P75e#5kAm3l zw&Atp-oU9)k8LT=x7tuUsb^=KUixf#D0yIZy^|8ls&sF7T{zM!7=)&|X# zeO=+J{j#zfwP2wc*$fVSVRb@fBZ5E&*20Ij*lO3%e)g-iUQMHM^>ejTUcfjv-v6B5v?kH$Bhw zrnMs0eQ~%~EAi@Xg#e+pV=Svtq`JMxEBx7tyyYQUwMV!yygVciDbBEo!2^>#_uG8G zCt?eGum^4`opa%q+x1l1U13{QNanGKvvpW{fLZ5R*BQCTIX%do-!nVgx2b$NWMNIA zY0u~O7|nDf+ZXi8ySKFQ*j|!iBM)}4xSZBE`%qK7=Zm=yL7)`T8vAZJP5;BWr28EN z8}T~l=*^W)A=Aem7}Er}M0p1`l>`yjW?$kXRKAaoo}g9BkJ;jhN}!ClmEmj@#&Fea$hfl`dKz**-BKY$anH-&ImO z)70Z{d$E<$Q<{CU`OZW{_Qla1*-#O?T(Z1Z$K=H3iCa%xg-a;awxw^H@(a$1ollS6 zSS)Rk`Bpq#2NQ=DcM%ITWzC%pq6OaT*HqpA^;yo=vGtLmH=db&{K7wo6<%HYspf4? z>x2wK;9yh1u5x(#zwS&ha9rhIF-K=))FPPXL?NFjwL#y7jrk1`XU$Z*W;mfYH ziHCtBzDLYO?QkBtbrTyes`7s>KCrprrt*4zr=$pn=m>7l>FQJBO67$&to&`JE3)on zXHq{ZTnRe1#l&<+LO8T4D)KPbsef!)C9*oJ`l7xi=2aDk3Ep#%4baISRuR}>faE_x zMDfi!*9F&W2AEYJ*=N=;#kxQ9W8a}NZZ}g?|B@qVr&j!<%hIDQ+nSupZv8m$*v8Dv zeHfY0j+p5Iv&>E(om$hR(;Pm>leE#k_$n>?eD?M^zLoEeXjqz5+vav@X?H(wc}X+R z;IUKFjFtVuCvoOYu08HOZHG@UJ*3NPhcCDLh19*%QYU67R}h2y`W;9E-S?~wx~}P< z#16^j^Sll?pi@yX){rKon6p>ct{2NUxEZb2+fx{(5Glt!>x3cT{qmS$F<( zBBy({?=_)zvl?5&U0Fd%Svk+%W?wn+v_lw~K4H62XfBKxtr*$c66}Ba%FY=Plt#q^ zI7lpdxW#b6aEGvqL}ReTA2EdCN8!euki)-iaNtQQWZv!>5D z#Iy#_^}N=ex}jiD^Xx#_#k~G*cR{?4OKS`?lDDp{Dn>6%#voAmoOWcO%ExN^!=6|0 z0T*i(>4mouFcdQzdPZ>CHNa+C$IoMp~f8kU_VVHCbrM^S|i#%=>ec64|1+`VUH+Du z^v6dAq$OYfBm7WMZ%=BCu`O9N0he^_$m?mpvQsmvUR`c(Pwn#3qF*MDbPG67<^?~z zt?ytf-Tda5n@ZKF1##U7A-FZkNXBBt@Tfkz%54GlkPH&-Cxr?)r2|PV)Ws{s=5pdeWoi)%~{T z33HF**o=-}@DO|V^rBC%qGKIdMbFvQlv?T;v%_@Od~Hv4(={%h4_X1?>fHfvKg7jaf8^{mN5cJM z@|GW1z8*2ME1omGOo^F2Duk9S$X(y)cb`NOEF3qALwu0_ut|S%e8MouAm87kw za*L-n5Faa&@N|mn>7d+9&f5u57kF*ya%ujWJ;{@#XB)hBhOYf!SSo!g2RF)lYkGh3 z-e7-zvx{ZJmUqPL?wse}@XsZ$K&|`jn|ZISj*(8KiCKIu-dLid=$It={Qc3voD{d< zAd&$uDlE9P-Q)IhLU&TF`?{7)BT>D0Yh3yzRRjF=(b?yf*Jla)j=x*?DSZEvc{%Pw zeakhDEHj#OSiwYj@J48<^00s58t@hdJ~vTZwO<9k)i3LvtdnleiTpU{Ahcs+3uP-ec2ixNRcC%n z5?7lZL0yd((=W-Rf~sBN-|zM)`nuS&atYEM{nVa5zW4WJ*=~_5pS%v`^Dg;9`&!EB z{q+&QX~4lZ1~2)eMp8@bDK7+)nr&97TT~mN4<)?u%9YNo+j2V0OJh%zZrT?2r_l`! zTe`W&t0&x_?-28F7&_RZCoYTYvDYZ_43&6YXNl8D_i(z z+L>C4{9<=r%bw-QG%qbblS8ubE7CYOT6+4(%iZ=0O!UPBP+@SyKfj&o=Bn|!Vq9ay zaOCvn83{j8{SpPYsV5g$GiA(;4_9J0ck_m9<3cB2d?9&0=R(Mh811({&&R?Vp%NVu z)ETrUiavgpaDrb&Bn!LMNy#O?UJtz@>oxtX`bsI{*G9xK+D6K_)pqIVV)6Jl*m#+2 z*@LG81?4B&2s@<{)rX;HBvN#H!Rh2%TfJ{3BUIe(vdQyRKXv6dQSPWUAjJwLW8R3K z4G=$6s@8T!Wz^NAvs5DEVUO8B#wTIFs@;!VO;cLwqPIU~Rdj|C1!`trL=#q;#621v zpXK+IMUrYo>`oo_=Svw{KeC=L&Haeb6W96AE9WBZ#F=Q23GBlX{{q~olZ@am{YR4g?EbY=6=#n4PTjHI-Seec!)Un0(rEd@6zh4&(twm zmo~qW(Qfgv5X$7!;&mE1AyAlae%iM2iN)mDEQWoLZ)^dli}3qZ73|vH+ixy^qq$6X zJlXPmLcIHC{u=^A?#{xFBX=*eToAIFlyaxS8ac?j7aboHo%1_&yT9ZT~dV>m?SU;?p}5{c;nw8N=742?2?*)xS>tIv(7)z2X%%;wP^->&E@#MN_}V)@wY~OzdsqZ}gfv z)th@6i%XG}Ao{Q5h;%F|o1qj}y;)|45F{)>=Rj;b_AnWJzhG76_eLMahQ zC^X(;J9>Ees*v<*dsKrks$+ej{7U&4M=e)uy_W=;Tqw%Jnb*p5fv+L7&CRcXWtqiw zA;I4{!ZLx-AlCmmvd1;>)fq8Cssf5A{j`x{nK@oY(xa?gC!$O4=6?xj7Ak%}{J{CbnLK->ia6ix#1|!+62U5$ zTazj)vES*tHjRr-@z2Z3j`%pSC>Bp@ekN^td;j=Di#M6M$7SXAitc%2eawN^a6m)x z9L99$8mry2Ybhn_kEXVY2B^$15ta!`#{Y$^SOV4(7)hBW3PZ3!5pc_$`C;Im!vS`D z?ED=p<(kFRt2wu5j^F}s^w`&PW0=hTD^C0Y_5K1FU@=v|L@(i3_#yk4OIO$k?$yqgS|VO{7vFmoA~TL;h4aAE7aW;q9s8&V#;AP4d0-|t#v72G4s#G2; z#dXH=D{n8qo9c%U&u&Orv8>vCwKS7ynL!b}@}E_rcfyFu265=lW`{}~{Wh$YWJp0* z&l6U6x)K7?3cnvSlsL;LKeF}!>QMX49oNd&`-xVn8`;+HwDpON)z^adds_My|n(x z^OM}O(U6%md0X=Y^l*Bmbljno5;?z9ULWrop}5(;YxAcM$7qgmg|yruzaQosfAJz< zB*fA80PpsPJZ&e$){BG&oU)&P-uwC5{Jsd)t}8+pgr}bGPwx$A-R+zxKx2s+lqnq9 zt7z9Jr1Cfk+d6#qYeZ+YV_I#P@zm$7ypCni{y4)}_gjrvgelJpej|~)7%O86GQE2pNXk?}25c|}5Hg}hAuL(YQzF{{1 zqgMGZJd#O=s+&Di=x;p;G&#=m{334LLCI%?a(Gk(9#7;AQ_oIze%A5(u=~Re=ONElX}l+{BfU6P#bM!Y=~{c7#((1X-+9bB3R@h~S=ws@mRj}al$PnN$X zXRrzTWwke*)hbf7Y$w#6@7PxfQCrun=W58{j}hlGG+8l!fqMM(u9%G)O>cG9o=D_= zg`cOSY z&sFImb$gLj88SZDCY=@8T=EeP`$FEQqB!lhna7_RKjZY^`Wv>m4`)g5rq(29kC#4= z7-Lzky}|8w)^JSXG&s}zv!$Atw|j%t`&x^@qv}V6YeiSas9pYkMCaD3#5511=eaY< zJ5z5_W+jNe?KLlTB27~~XMDQvTr8VA!?iqM-rSPsykwc-ZTovZWjF)m!zSgpw7pm2 zM~V&}vpo|y{i44|Cgim`ropZJ8WB~Z-pQR{dbQ=^n;zAa&8*Qi3GcKE5M!eECG(EU z4WyG!nJTjB8DPj;tN!$MOr#lj_z`SVosSJW_dqGYpMi*7#&g?1wKaiofkr zVxl)@$N`V~v)`>0-Qg*(^Pu+I2FpI&s_S9Y$j;?EIhZeLgbjGZk=% z#jzR1IIaHg0XL2bRzTzbW?Oo^V>J*@XDTz4ldZe!vYOdHIg#}Fy^`nHxzbVX(hmj0 zPQ>aMJC3KjJd?CEJFl@>qN>J#$Wvsl#AfDU%!SR$H+q`#`Z`6eDv)^oKKrSJwjo<@ zoUYIvwTLt4#sy5jAZba3DMA$Pwo>wtY1^$(f%J9vsZYADo~P#Wi&n}t<*7DDQBMka ze)>L@ujH_MFc8J$03XK$jr@&xSVF}DU|yN1AEB9gB3k<&v31**-MsxWbk(b)wktF@ zWWRC>F8=xk}nm4+|#MJs!^ z<-J~x6Rr_7xfQh5L*sRser`?o-cfnOMBN6#L=p3IE>Wsmwi|DEN!(s38YYoxQ|^$S zVJTZqdRF1t+<4=m-GesWIKQ}W(XUrnelvSxpFC|6b}NzrfK!2FL~Vw+251O} z9PFo+!)mYTHG$XWD-)CS^tm11qhCfu8wT4#RcCr{R>j@c41ad9&)a^)x&cv^0Z(YlVxE7^RVEH#~-l?9_;_w!N-j*nX8%w(`KFTii=~egZ)4M zH34Ik;|m$2?LDLG_az4knud37nw2cmetdi)^ERn@Reuk$*tU0RIZmbJ4SQfVD(@N0-EMz{S*;4C_A+d&T?x%U~(hYSJ)5Z>!X}^!CPgK*`hHgJP zPKfr-W0x8cEs(r@bFIpy*5!3>nGbf1IyQ<%%EVf8mKTb8&Fth)$rm~QLg*ptaR-X* zocrt&`~6qWl+ekSFYp%|qb4kp9Gd(lQ}68Oxa%>uQSjDN7lFj1+s>8ot@c!|kE+k! zrQE#h9Y2w~KFZ}(raYI6ht#>C_~`oX?D7JqmR3=_5AW7kz7CFLvJlKb;$PqWX_=PYG-i6RLb$Rg@a9XRBcuM6z1b{xq_*r9*5mM}9igKsG|mJSKg^2?dHz z)zJaM>dTi?$OZA2V$PtBf2{Bl5Z#@(Q`|psYuHU$_ZHl`B@qXwvPT=fs|dv?lMu{+dXbmIy0Os>ab5ydL6P<`27b0>pBb1 z(izR|_1}Hge4t%UdJpVURY_!@h0WaRTV*4>Uf zqkh*RKqkK(GozLM!WWw)7;x)=4asa)V&U)=-{;OT@t@a#z zu4$XD*F)kC$3I)WvrT&6Ri!#~yr#l!w|ZWE^f*WJB^K+W-)mY!j2b`h>KRUxm2}~c zoC$d&P9mu0@?5bTrxyQHK6FyfPvHHxyE>B->{lezo-h%c2})D?i=$5L6&wj1c@V%S zZa00zS*cgQeByRRxV`ifvdxsSu~{Z0<@U;4EImS4(S=OQVbRoUIqSHgcq+O)vnlDv&8hs+AeEH;{kX&-TzMkx>gpQ@yU2a>=Su22KC@rjuQ^~U z5V^sQe0)}5PIGfCGSDn<%bmF!cVh2ai8k0rR3wXK%#e}_#a@y%f-)4dY6kmm@4=0v zRw&SpUahP{o=>cgc8dyY@JzqTlBvd`(|5PfLT8<5_x%F$mP>KabFX-JRosYi(_NI^1CWhL`1ud!TLsqa@DGT?F_D!d#n@QcJ3t;Y- z;5*i@;@1;b`y7Av0sojC?jDBkehY8u9@aY5?`3mbV zb1bK~S*HH3ACTSf(?KkJ#%1I|BDI}pl*O8fg@vw_oGxXLla)W&$%h8V6Dp0adDA=R79cRw;lZdTXR4lVdl(18YW0{WF%!Gka|8KYY|S;{Fxm3qX~i3^>lNi zfD}tOr*bC*f4HRiIk|vzPavNkoO|g-aicF*sh;j0bO;CBiN^sy4~h?Hrcgb7T_GF@ z;z1OuFNE~)BLjD`B~LGq^|{IZmcrTt&l z4VZ+NlRFJ)?ypu*i{=8Svl21D0m3uDHwqXmkXpydZ~KCNPza>Wk19x(1!=sVE?PeB zUKGGMl1!&~Tfr3*r$B~8Kp?Pp{h8RykPjATCqzJ@{(lt2|LZtda43cYHi`qAJ!Pyi zM8FadL=c#jlu-}{tUZ-n<9IN+5?p|lVu*MIC=dY{C>{?60jHKqz@U|s0Rt#dhp+H+ zEGQsNSqb!q$0ULnMo?a8gRcx>uu2H{3Nrx<1s4KX5V$WM2)-}|pat_X>R@w-0k#hp zyb_UdB?6E_0PO@c1_-z?FAmI$2eynQ!u@a%9$b}yvfwMDU|>1|LkaK|o}aO3umS-O zK%oeDPCV$tP|6SH;48y^;F)mn4GLBf%r-076%{_h|rG$8^VGOz;^K8!q7Fa3)m)pDGYmnZQ@5^*vg_!!1nNKf%#(GgE3&+ z_)%cV2`D^RL%=I+156508LSCx7a$n@T7KBXk1MPcu1PBOVNf>k>8i99aq3#E)XR-EgSOXxBLB|Mm zi_wiSE#lwEwm3Hx0trzJS{fVx7iu7?;TC6K44R8lgCNP^U?CKOWDN$rKd{qcN0ibB@5TZXKcB<&Kb~u$e~tD3{)7XmqXmeN zE=~{{9vG>G`CrLE-26#~{a?vYiXd|RB*gt6gdl1ILX@HsnEDs7U@;ci{wd2ZR{FCM z*6j}g_0bAHQJI1hmFl~HVfldT`LEhw?f&Ss5Zo7j4}&!FbcGQQcuV09j30~mgJ8_i z%wp0r{ukzuI^YK%Fs|DI0GOYbFWt-284AEEqEU)idB}rK_oJyGk>3A4Q}m^}%OH>d ziMjf_fFF?j+Re|^4RUsJ*$bxmF$5rGOyKG3qYMAqW{a*027G-+DPe$1MPczcIh2w# z3MDP`i|PWY6gLELp9BOTS^NWG3=iH7H|XCs7zM(Akk4Wp8jAw>YDpXLG{4bd;1IKf z4$HutOUHw~!qPVQ!1+II;D!Jcx}*(0A6?RhCII)nqz#2ZE#(CT@)nl1q0tx+oBrDu zjYfe_PfOZRc*1WyW3a#Tj3$6E|KGkSv@+}g;J<%$ib4}X2w&2MLBn3~zjPQJ{NgTY z!{dLO14D%U)_>`+=-=c;VR66f6br+gC4F%y04tWZVTnuE9EDQ`$skMU@F+OZAO8C( zFCMM*TVFg5#-2;)@W0D}C;YY^C_Eg-mi8qAIJTsXfKmR9j(~?_(SPS9DF3c!0`Ygf zC}BVdUeZ?ykNRzml$7A8_1`&^lwsff-!^3wz)efra3BOPZBu4KM+A^`2^|rw3>*~v z_fz)(LjO(&L)oQt%EYBS14IR|_5@J>Sf8pL^f)-h0ov=bZPPbMJdz2}4b7Ig~sWCQ FQ|f{ z;7GWK{b87r5*%Ub=ShMiH0-E$ZXV8XgrS`?i2_H14#sd56&Q)^2z6BWV}P~?nF`0y zIv~vV*&il3P~n87OD$jOZWF2y&Xf?F~DI{Mi9HHwDNKjv{HI{2#xb+|8p#EL}L^nje7YR_eL_WfpMDg(Uascar z=m+@qNsg{|>K?uTBNF^E7?eC3j>cl;6~Oce4X`jUngS;-k*7u`dw?qx)b&qhK(~K! zZ!d}LOm%^ymN9C(x`Fk=5!!Bm37RAa4@VM2Hiha%vU7*|rrkQwPTox5AO0+JyDT8} z<*SR4@{@Sn$n_z%iHsKY3p;M|Dcko)^qe}$6P76ygWbg8$F?!BM&lX9_fQSy$H|Xh z(k&OfDV?W}x91EO`r9|`dTHU+y6dIiiw2JmvMCL2on_n8k(M}dQWV>TRv&8DQ&#+) zx+KZMYP%*1>5+j%h;IFpQ{jtcqH->wU!Lv0T(P6LA8L4dU!V1 zOGb&U#@%(!8GaLY46{D!C6JFw{p{b8a?{|pt^26GKw>v+``Hv1RI@sQR3;rUXq`R7 zd9&`_4+M*gN87il;QMSGEgOe77F`O3M`i8!u7b#Zu9y8qV6RwaxAwt%1YJT{S7GVM zH4{fs}RJ2&V-yw3l@LaiL&H?H1FT6O0FAQ!r#z~T`=}c#( zB7$448=2cY75^!5FR^>%(v7Cg5AV0PL@=w%__i0|sxOzk+3Lx0nb*Rr0|J1f9#8gaB`ej zMr{Tbd;5|t|AzgsQPFBo_?tbSz{y4fn@4O)wyJH9h^tG{ec9m8Z^@%^L{dHX{jIR6 zZKK_m975~%m*BfM&Zzcoa|z?Gz)dStC56I#wq>1bn=kX{AK=#5EfJJYm`>l>sz42A zw0oaXa0#>aq}Z9aP7YTz^~X`EyKP$f_*#=~V`cMh?`_)=p7xH7G;}e&Sn}%=CAYis z_%5D7t}6$gBoEhot9YD0=T2^advnlZaNuFLx5=b?AGL@lRla7kOUEWxhpU;wk7O_OGBjFjuOqVH)A+s?&Bre~gO+4g&M7PJ3kIgV5f8Uh~o{Qs$=1xZr2VWa`H;@|no>})5%T&+n zW`z8esfw-zzxD&4hl7P{tmJ)8aACGp6=n{eNUM4H{q^*f6B&KqMt0N`=`p7v+zfnY z)=Kfjj1^o?E4ylDGa>b=ld!*dRx#c$PdHt?<=n|8^4)pti%rMOhxx6yT}+PYOj4Y{ z>a}n2xirjfT>AD$i3H5=Id4MnSEi&Lubw+aFnV?HPFLKaRyr&)CG&0Khs^1+Em$_J zG{XTI6BH7S`(?c z>3edzRWJO`6lXl`b{52GI5bDX$1^sxR77f>k<#_ua!ozVSMh75)hX9oIG^Hbd1`J4 zEOM9lNZ@y-?wy-t=`7+;_podcF-XFH6Q`V@M`l8CYRsyH$givY)_?hP){d$R&rT0Y zO1wF{rC(6%a8lJ?3lotz?7866Z$5aIT>7-rt;@-2#4;l};#J~!w}Aar#)yZEb^uyx| zQRtDwd#&A!3yyD9veDc>F~Pb!Ld+p~!}zEg`6aT=F=nUU2`{NBxZD;hGM)Vr(Bv|KW=#$_w!ja zvtI0V7tzln#pGUj+Zq!^Eqg};uR_mSTDm@ZQ4@1EUi*+<<@#lTXH%6Zzb#?!AkK(D%@M{Uc_iZg{T=3(2zS{@gU(9{+nMIe3>EA5J@=Djl*L1BhhF>Bo=L?gBk%5}}r3|?nCwEP*zEI&AC4C=p$ z!R;d^FpF}w`-P|=_ir3jy?58gYjgSZc(t&jMGQ+!jz#cwz@4_Z3HokmRd zHs1SM#>1CCJ0c>M7|CH58t>U#I7+@8vtiNae3v=vG=A=^h0VIBb$fdc2s+CTvN*V= zH4t0LQPzAV(cVSZOw4y9(>a5w<@JehifPtNYYh7JghIUJ<;dx;gR(xvQ{3-3@X?|I z<8#vfLHIKnlYY~To4nK(t%v7Cw5lg0;G5Y6dyVX;(q080y5mC?xwV6rqMoWI&nMdX zMmlF*s$pSi=2c@O&pefD6NOh!zdB?kM~7cfn6*N7Xe{x+`K4^QGe%RC^Bpk+@nPF1 zaZeHL0y(GY=hvB1r409;D#vW=ULUxd6O~vzDp7Rxdf>fC^^fG2@6Oc03p9-35Bzz; zsgpVQcs@npG|VnLfvrPzI#@7eGilOHM=~N`EaojHM(VQk@hiT9 zvhQ2)2PG4f$KWrFjU!sKuOznZa&Jq7DLVBq%kox^IP&Q$v{&mIM+qgO--_h;h@C7{ zYROcbaMXWLD4x>)Xy0JU+!oJ@Lj#Tm=bNb_9dl`A56%z;s%A$c@T~gL0~3?8e6G?6 z<7#2cOQ*bf&kt`L-^iQnd|K$a-Gy}wAzhr zucg$R$VNh`ylU(1#^VKYvkb3T)IB$vdN+$^IqVUY&Ec?RhP{kM-Cd`F@vZ6_&B;X5 zqPQ66ZG86xhMn!V*pBzyWVkM5s+n|rvN=5OL2guyW`oK;{ck&x_R5(xkxhgSEhI)g zK0#%X&cl96{M5a_w41{`!jJbWaY1Kpyw^=MRFT~KDdN>uOe322fvq0VHa&TC_sF1y z_xkI?mA4;I1Gk><_1s(b?o`W{eLty#mYv=;UThhvuk_VERC8a~E@@~$t6Ykf^3RaH zu_?7i8*|@o-TL0y`WB{5KfG$zVXqw#D>E0nc*|Zh`m7s8`3`gM-swhSZ#`d~+w`U0 z^qUxrv}n*kRp{Y&sy?N6KYNy6*@c5=N>KNIp}$kCLj;zE-YF+@>v+Rka^SaJ?ZzK|w9d{67xQ#3 zSf$b7AhbnZ!9n6GWU3W?e|;^Sjd{^jjU5IX*th*0aO%3op6WAr>A_$(}FOX_mv` zySNe^D(M8J80;$imlm)lnkcrxdnTTNIxmvrK9}V~ao@p517%6WUN8x@gWbm_*FWbGz8)kKWOw6; z{F(P8^y9^EVS-lECui-OgYB%Z!T9Jn5y`hnfM$`MghcX{CTxcWJ^V?}GYZ5->13Nmw(-oH@_ z`s%FZr5*J+QbAEP!SW$7=;gz!B9^08gp!x;;w6;_o=Cg$w zQTWv;QC#y;77&(C=QBzm17KeYn?Lis`&k@1^GMBqqdKSkzPcM$2aeiN>tieAFkGj1 z3r#hBLp#(wF^5Y|Fmc956!c_$^Jx^ypBZ~%e?2q93ZW>*+mSF@pelYs@n-Y6@^Z`% zs?F9((P_R#Y3XsY9fN%Sl4`;FeFrpHRl91%G@a3I)v{hmP;`D?*Z>?-z(bXk?O&P;y8 zURl--#-1cknAqj}=hv>Obmo$zgO*O{1>C)c^`|!le@V|u&0HTEp8)i?)4ibwJ#P; zanDA;_eq-UZhQ~#|G+94ee!&PjOQh{uRVH5Zss5Ad>6(dRo`(2Hg%fJjPXv6j{1xT z+Isk{*Xdu^5-++@ILPOc)#A(E`8$h8Lw9!F61u)+`sJ|;y*|x{>=Ohi43R@px#LIV zE&GHNpPj=rkL7$1eNbteTzzKm^!%>gu5Uk?M5LU)(P1T>cmsuBg?4Gh?c`+=<DP4nI@P7xZ`_C}JWgtXyK4Wx9yj*5zST-QxrzAQ@;X3gZr!Vdt$kTaLvvt1>I@?O#yH(o4?o+mg z_t~5imz{DJxWB}WUJ&UQtupCNI=)dbDZ+1$a7n7{tS1Nl?wV9y3dhG;whuA^Mzxw( z$scEVsw{X_H}Nm_BC6*HyJ=a~LJ(V=;!@sfKJb!;J8- z;acrrm0Ml4@=WEr)yEw6%%?d@-DSqZzUG<;w|%gm>@f3qy3jn)=49j+cm7My8}7ED zt92@RXihWH=;r+>wqb=!8UQ)GD4}Wy8v7W)x?;b;u_3TL1TX&VrG8 zIg2wo^Iye!1htH%aVZH>2=7D*3zKJxCz2Yt{X$=n9q|&LRuh=U#k(*WOKWgyDCvd;UEt>!MMByqniIt2ed)$3#la|NLw0=uMYK^$H zd5JvldRLs9epjufTtIb&&Cktc;#HigUn3fhrW(hj-#g-&-Gz81Y4Vf~i|E8-t1W?CG$2NG`$ZBKqqUqLYu@iW$#49&C~6`(gCyBdNeQ%INz}rFV$P0_6wXaR#@W zir+rkd43yHL{;2p^=#NXk;f7lu`+`fj4v6;z1o<3W#agWm^I3c&!UHVn%g%-Y2T8| z?Kym2fd|QF(sTZ3L405$-&zgl-3PE7*)47U>bSaSVtUG*dQV%X_R`?2>>}eD9umU@ zqxFE|D0WX!%$mCAJEFdPTbsXNL!iSqdU4GvBJAbkDlC#UK|546DLz8|tmp=vH4PoF zg4kaNSgcjul%c?SeS4V985yoMFUq(s{EZq*C#XT;k^dbIPepWj90O^J4TRriqz**yuG_XQJ#4X-(b@7$rYSH~ei zr(EQWc&d4+^@S8u=~CktWv-2N_xdfLv}i_qMlVFXS!=qm@2yqh2mLd}38-%yO>1_Z ziPD|3(HgQuz5l6+ED`$3?xD4vB{-raum5hBZD_|VtWTl&GuPfA!aF((r4ub-Fsle2 z<&5>b8sVIS6qxl@%pB{O(i>`(@14c+gmYTj+_p2QSD_JmPFTPzGJ6{;qB~SWUKAH} zgmmc)wkU?T$3@;wkPA7kDLaO0P-+O8F8C-SGGExyTFTml=s(6I^Yw8sNp!SPXuwXO z*-l{0{D!w^{&8QEtopodPe&scRT?RB%yK8Ry8{l5T`7(knNJXOoxR}x=(_)ny*u-> zZftSeN3q#{$kMq|QKXBt_3QPg>y4`!`9mr(4=68oJTlWR6ZFCB9JX@vi1}f@r#pmf zrF~XB?n<7R($%cij_t6Qv1uY+(oA-5Q?FCKHt#=r?Q#e1Vq^tI*Xi9|osyI~KWl~3 zpPVL6xBMi#yfPl=u}zma>veY%uV22CWZ^=+$XAmU+|^{T&*1I>%sWw{0DGRj<$?RI z{3-VAgS>}Kj7(Dfg?@hjH1%=kshVxbc$5>}yp7!2M71l34J-kh9J*lX=1 zT2QU>E1>U??Ff_==ZH(%BCo;WfONNa%d&N{U+ZWAN5QM05dT6d9}^vF=md}$aG3nx zZ(zV)mj&!~yQFTbW`EfC;B$U{8!^$jK%8rvFm zuUhNQUro8NM>8S$-N_R5nb6t ziG)c-9pEU<6>F z9%Jiq`MFuJZNxz!Az=5l5?(G>h1&4i%Qgy)HlO*3+_m8jmr`Xp9bC3w^N)$B?Y>-^ zZP(N+V)^AWkLjBeVRQsRFZ^6Z7(la#0rmVE*RgAaUo#ognVC+ohr+BTu#MH@W)=l| zgtUy9+UBj}N6~LWD2wj{E|K{hltWAdD}&ON<&Q9364)v#RQ|CdFDXXLLdDeI*o>!j zB&@KhU3BQG?1vy-=}=L_$P1g|w}6E9(aO<#q=63M{hzWV_?2@fC-PWOY*b>SXY zUOzw6QL1s`PD@jt_OKAsTa(oJvhxGHgt23!kOubF0^aw=p*p!WW};7~Q_#f~IA`V- zZbmzNa3Y7veVF&7X|*Q`Z1HcV_vdN6gv>ENj(*VkMWGif>|Iv(}JWphUt2Vq+C`=J{NA$d9@6V)vGnfr3R5c{5Q zO8oa?CHRCBTb^@srJp_|k{OkTf5!MaW0$)CEBDLNdYj5Xfjt|JB+N48?VSjGvsnD` z*4@6ot-|jEn&=3JPNas$t-`O%8R>d4sj*de$;<>hU3>13EE6`i*vHn!x|Sx(!b3L z8xVzCX1?H=HW;}Uep1z`{XypWLb-YFxv$>%1Hnh5$2(g;)tTBhUhE^Q;-9>HUK2k0 zJ*s$oe13uHknO>6Iy|HkQleIK@U0)G%dG+Sp|fdq8*+m&wc|?n@~d2??-gmVdIv6y z!9A6fx{q#?r#enV#mHRTGFn>-cX-*Ou-JNAr2(ZjoFQkP?9<%e!Aq$_qZrGzTl<@P&N}d@E**-A7a4@&9 z#^9t2zlJx*NgI~4{q=9RNOEz6Ilb}}vPhLOH(w*P-RZTV=!MWN@(w1Hs|>1IO*yuk z@^>vois$2ar~f7iM=w%b%{j1g?0xQfB%8+cJZw=l>#0=wiG9}&hfOUOi|bW$_%k|0 zg>-js&QGlmwY@mn_)x<1$Q*y8ssy1W#g;oe>!I-Ocz=6^jaDLMn`2&07*3P}4ro za_5-*Pdclj6BZKi|7|~&qjqu!AZhBg+vU_YR`u|}XQ#8J$|8TvY#VIsv9*k6tf*}r zQ^|Ye)wONsiFFr6#|~_^ymh_pVJc^2YScZh=fnDa961UbYU$8KCniFpRv~?y5s@35 zy-dyu{8Zf*h49^%v7__nz0Rl}Gm-mNp=F7pDW8nb<%+&CQSncaPpca0>o|-ZPb!n6 z1l%sKK@=s_MmU9^x$k=6HbbfsgGOIZu93zDk?zOYCOd8nF{_-dAoC16+zL!C7afm) z+wTyPRqM>t;uCc5lFY*016$u`$?u0-OtPDjHU_R`FjLslH_`Xuyz;*7JyZeo;{sg! zMs_|eF{SyKoF{zmEU^#KypNmK7anG++4%!$ux1V4(RN11D|-@4#~M6JW=hu{`1X_G z%5Kx7zxe}L@y~mcrPM}Rz9K^1jzU^)=^Ghon#pLA4%?Y|o7j;lkVL&GR1Ft9FEALf z*KWCu#^GQHXpAO_;^5`#N%io8qm~mJO}y=?w8TjW2Z#_1?A$^6q58_Hm&U6vr5wtE z93K>14hxRTPz3NY2ZaW)1;{&uQ0UuHybV-noNK2lC zA+*RKyAVpI{3nU?7YK&XbaiqffmBE+_wgVcN(Z5M+Btxsnsx-O$ zdSCXB-2adqjm`fplkb zg$gvsuYv=)FeoI%0Fdg2R{#?c72pK02gC!F#sTC6sNg^jFA)a=1&+tz;W!+a1UzE| zpwa{dKmbrYS`FRDfC93I31B=lCsBb2>LnOd(-wmvz@Qo;0$LPQ@Ic|9u{f~sr8z)f zfS1+=Ck1GN05Rwj3Gkfuy7ox6#%oKnpV(IKRk^I zPz~Xy2@PT(paC!x2I0hkF*K(9A`Yr)`U7FYLKjGgMG)E0(x4g?nmizMO9h5omf7#A zmKuP-r76&0#b6i?q6$zB4Tsd=m$-h_KxE)xZi!2@8L$M<38Xbb5(%T-Tb_W3C&GUfND-EF0Mdhh7aFgDUO+nWN1^Ehq!YgiO;?t60@8;+ z3q%*~8q5Ld#;*dc9FN2S83I}%9iUT)3P2{1UI5?oXZfWQziNmnK<+SF4YmI$G+8g* zhx+_3kX}Gih3bD&LI2#?=>cB?EQv8>ZUGZ%EeALXiV*>9$=1_S-oTY397vZ2C&9rx z4A6ot8*w1)B{P>CPg>^DMJ$@)X@L48}OL_bi@45Zb*6HgF#2AR#EwbT8O%3SNac$f9TvlK!`bi z^@JAxZ%@4RO$kEJ)e(wDz#|25hxWmRH}IUa^(?1DgX-pL z5BI^yqmc3#S-1<8>Pb;VAl!f7llSm)mVzNbl;r5`06xO_bC{>26Wrd;;RrzUYYNbn zhQQT>tOaK_h*j|yhIf8FS!x1kdIo?4^Q|D2XOH#WeLKXzjPM`)Loor$YI3D^=i!q~$DGJ`ayqDyuH4rIws0Ygin+IdlFW(S2qp%q{f5?cGTVgCn5 CYYIgG literal 0 HcmV?d00001 diff --git a/paper/tex/marc_aaai.tex b/paper/tex/marc_aaai.tex index 52e4d03..77841d2 100644 --- a/paper/tex/marc_aaai.tex +++ b/paper/tex/marc_aaai.tex @@ -22,6 +22,11 @@ \title{Learn the Structure, Not the Values:\\ A Controlled Characterization of Learning in Exact Constraint Solving} +% Submission is double-blind: keep the anonymous line active until camera-ready. +% Camera-ready author block (equal contribution), swap in after acceptance: +% \author{Quang Bui\equalcontrib, Sparsh Roy\equalcontrib, +% Akash Gundimeda\equalcontrib, Davin Yin\equalcontrib} +% \affiliations{SAID Laboratory} \author{Anonymous submission} \date{} @@ -66,14 +71,14 @@ \section{Introduction} has been owned by numerical analysis for sixty years. The discrete decision asks which representation makes the problem tractable at all: which auxiliary variable, substitution, or defining relation to introduce. Classical numerical solvers take the representation as given -and make this choice only by enumeration, and for the value decision they are very hard to beat. This paper measures both decisions under -controlled conditions. The short answer is that learning earns its place on the discrete one, -where enumeration is the classical fallback; on the value decision it helps only in a narrow -regime, and we can say exactly which one. - -The measurement matters because the usual comparison is the wrong one: a learned proposal -plus refinement is compared against a cold start and wins, conflating the value of diverse -initializations with that of the learned model. The control that separates them is random +and make this choice only by enumeration; for the value decision they are very hard to +beat. This paper measures both decisions under controlled conditions: learning earns its +place on the discrete one, where enumeration is the classical fallback, and helps on the +value decision only in a narrow regime we can name exactly. + +The measurement matters because the usual comparison is the wrong one: learned proposal +plus refinement beats a cold start, conflating the value of diverse initializations with +that of the learned model. The control that separates them is random multi-start at the same polish and budget; adding it shrank our own headline claim, and following where the learned component does pay is this paper's subject. @@ -95,14 +100,13 @@ \section{Introduction} shrunken claim compresses into a law --- a parameter-free factorization in the measured single-start reachability $q(n)$ reproduces the separable family's best-of-$K$ curve and predicts where the favorable regime can occur at all. -Fourth, the same discipline turned -on our own pilot exposes a trap: defining ``failure'' on a single stochastic stream -manufactures repair effects that two-stream selection and a budget-held screen dissolve, and we state the protocol in -reusable form. Alongside the arc we contribute the substrate itself, a pre-registered +Fourth, the same discipline turned on our own pilot exposes a trap --- single-stream +``failure'' selection manufactures repair effects that two-stream selection and a +budget-held screen dissolve --- and we state the protocol in reusable form. Alongside the arc we contribute the substrate itself, a pre-registered entrapment result, a partial cross-family transfer result, and a MATH-benchmark scope -measurement showing autoformalization, not solving, is the binding constraint. We make no -claim against combinatorial solvers such as DIFUSCO; the problem class differs. Every rate -carries $N$ and a 95\% Wilson interval or $z$-test, and negatives are reported in full. +measurement showing autoformalization, not solving, is the binding constraint. We make no claim against combinatorial solvers such as DIFUSCO; the problem class +differs. Every rate carries $N$ and a 95\% Wilson interval or $z$-test; negatives are +reported in full. \section{Related Work} \label{sec:related} @@ -416,7 +420,7 @@ \subsection{A factorization law predicts both results} % coupled.json (R7), pointchain_learned.json (R25), real_systems.json (R26) \begin{figure}[ht] \centering -\includegraphics[width=\columnwidth]{../figures/fig_regime_map.pdf} +\includegraphics[width=0.88\columnwidth]{../figures/fig_regime_map.pdf} \caption{The regime map. Each measured family sits at its measured $\log q(n)$ slope (abscissa; the separable and coupled slopes are the fits of Table~\ref{tab:law}, the geometry slope that of Figure~\ref{fig:law}, the R27 slopes are inverted from the best-of-8 LM arm through Eq.~\eqref{eq:bestofk}) in its solution-structure @@ -527,9 +531,18 @@ \subsection{Relocating the learned component: structural repair beats its contro \end{tabular} \end{table} -Table~\ref{tab:repair} reports the three generalization tests. -The protocol (Data -Version 8) is deliberately strict: one reference solver certifies the data, grades every arm, +% provenance: paper/figures/fig_repair_accuracy.pdf (scripts/plot_repair.py --panel left) +\begin{figure}[t] +\centering +\includegraphics[width=0.8\columnwidth]{../figures/fig_repair_accuracy.pdf} +\caption{Table~\ref{tab:repair}, drawn: the ranker against its candidate-only and +random controls; the dotted line is $K{=}4$ chance. Menu-size scaling is in Appendix +Figure~\ref{fig:repair}.} +\label{fig:repair-acc} +\end{figure} + +Table~\ref{tab:repair} and Figure~\ref{fig:repair-acc} report the three generalization +tests under a deliberately strict protocol (Data Version 8): one reference solver certifies the data, grades every arm, and runs the end-to-end solves; nonlinear ``exactly one solvable option'' is an exact CAS theorem (a distractor proven to have no real solution is unsolvable at any budget); gold and distractor parameters share one support and one prior, so whatever surface-form signal @@ -566,8 +579,7 @@ \subsection{Relocating the learned component: structural repair beats its contro the ranker's single call solves $0.939$ ($N{=}360$) --- learning beats probing on accuracy and cost at once, since provably rootless distractors cannot be solved at any budget while short probes miss the gold. On linear menus the probe saturates and enumeration is -already perfect at $2.5$ calls, so the linear rows are mechanism evidence, not a deployment -case; the $K = 4$ checkpoint transfers to larger menus without retraining but its advantage is +already perfect at $2.5$ calls; the $K = 4$ checkpoint transfers to larger menus without retraining but its advantage is gone by $K = 16$, and direct $K = 16$ training sits at chance (both negatives kept in Appendix~\ref{app:kscaling}). ``Exactly one solvable option'' is an exact certificate for all linear menus (rank) and $99\%$ of nonlinear test menus (CAS real-root nonexistence), the @@ -581,8 +593,8 @@ \subsection{Relocating the learned component: structural repair beats its contro repair the failures the classical reference cannot. Two produce a two-stream failure population: far-side GPS trilateration ($0.848 \pm 0.020$ over three independent seed bases, $N{=}509$ failures) and a ghost-root conic--line intersection ($0.263 \pm 0.015$, $N{=}158$); -on 3R inverse kinematics and far circles the classical solver never fails, so both are -reported as negatives rather than averaged in as zeros. On +on 3R inverse kinematics and far circles the classical solver never fails --- reported +as negatives, not averaged in as zeros. On the two that bite, one construction chosen on a disjoint half of each failure pool repairs \emph{every} held-out instance ($1.000 \pm 0.000$ across seeds; pooled Wilson $[0.99, 1.00]$ and $[0.98, 1.00]$) against $0.433 \pm 0.049$ and $0.114 \pm 0.011$ for a restart control @@ -594,11 +606,9 @@ \subsection{Relocating the learned component: structural repair beats its contro design is the value of the structural decision, not the ranker. \paragraph{Where the anchor stops: a closed negative.} -That anchor holds where the failure is a systematic attractor one construction deletes -outright. Where failures are stochastic instead it does not, and the boundary is worth -measuring. We took the relocation thesis to the pruned point chains --- the discrete-branch -setting of distance geometry --- asking whether one derived construction repairs instances -the reference pipeline fails. The +The anchor holds where the failure is a systematic attractor one construction deletes +outright; where failures are stochastic it does not, and the boundary is worth measuring. +On the pruned point chains --- the discrete-branch setting of distance geometry --- the population definition decides the answer. Under two-stream selection ($N{=}367$ failures at the training chain lengths, three optimization seeds) the enumeration ceiling is $0.692$ at $72.7$ restarts per instance, which plain restart scaling @@ -634,9 +644,9 @@ \section{Limitations} proposal's advantage requires per-variable-separable solutions and a regime where random restart collapses; on coupled systems it never significantly beats random restart at any dimension tested. Even on favorable families it gains nothing at $n = 1$ (random restart -at ceiling), and its single-seed run drops to 0.250 at $n = 6$ --- a three-seed rerun -($N{=}120$ per cell) reads $0.983 \pm 0.014$ there but is seed-unstable at $n = 4$ -($0.658 \pm 0.484$, one seed at $0.100$) --- so the useful window is the crossover region, +at ceiling), and single cells are seed-noisy (one seed reads 0.250 at $n = 6$ where three +read $0.983 \pm 0.014$; $0.658 \pm 0.484$ at $n = 4$), so the useful window is the +crossover region, not high dimension per se. Those separable families are also block-decomposable, so a classical solver that read off separability would scale linearly there too; our controls use joint starts, so the crossover is demonstrated against joint-start search, not block @@ -667,15 +677,14 @@ \section{Conclusion} The decision classical solvers make only by enumeration is not which values to try but which structure to add, and that is where learning earns its place: the repair ranker matches the enumeration ceiling at a fraction of the calls where ``exactly one solvable option'' is a -theorem, and on hardened variants of named real systems, one derived construction repairs failures the -full enumeration budget in restarts does not reach. On the value decision the same protocol +theorem, and on hardened variants of named real systems a derived construction repairs +failures the full restart budget does not reach. On the value decision the same protocol returns a characterization instead --- the learned proposal pays only where search is independent across variables and dimension defeats random restart, and the advantage -disappears under coupling. What survives is the substrate and that division of labor. The -lesson we generalize is about evaluation: populations -defined by stochastic failure are artifacts of the draw unless the protocol makes them stable ---- the discipline of the Protocol section costs one extra solve per instance, and results -that omit it read as upper bounds. +disappears under coupling; what survives is the substrate and that division of labor. +The generalizable lesson is about evaluation: populations defined by stochastic failure are +artifacts of the draw unless the protocol makes them stable --- that discipline costs one +extra solve per instance, and results without it read as upper bounds. % aaai2026.sty already sets \bibliographystyle{aaai2026}; a second one breaks bibtex \bibliography{refs} diff --git a/paper/tex/marc_aaai_appendix.tex b/paper/tex/marc_aaai_appendix.tex index cc596e9..96e7ed6 100644 --- a/paper/tex/marc_aaai_appendix.tex +++ b/paper/tex/marc_aaai_appendix.tex @@ -135,16 +135,16 @@ \section{Repair: Menu-Size Scaling and Cost Accounting} advantage shrinks with $K$ and is gone by $K{=}16$ ($0.300/0.187/0.113$ against random $0.227/0.120/0.107$ at $K{=}4/8/16$, $N{=}300/150/150$); what survives at large $K$ is only the cost gap, as the enumeration it displaces grows to $9.05$ solver calls per instance -(Figure~\ref{fig:repair}, right). Directly training at $K{=}16$ performs at chance, so the +(Figure~\ref{fig:repair}). Directly training at $K{=}16$ performs at chance, so the transferred checkpoint is selected and both negatives are kept in the record. -% provenance: paper/figures/fig_repair.pdf (scripts/plot_repair.py), RESULTS.md R10 +% provenance: paper/figures/fig_repair_kscaling.pdf (scripts/plot_repair.py --panel right), RESULTS.md R10 \begin{figure}[ht] \centering -\includegraphics[width=0.85\textwidth]{../figures/fig_repair.pdf} -\caption{Structural repair, measured. Left: the operator-aware ranker against its -candidate-only and random controls on the three generalization tests of Table~\ref{tab:repair}; the dotted line is $K{=}4$ chance. Right: the $K{=}4$ checkpoint evaluated zero-shot at -larger menus retains an accuracy edge at $K{=}8$ that closes by $K{=}16$, while the blind -enumeration it displaces grows from $2.5$ to $9.1$ solver calls per instance.} +\includegraphics[width=0.6\textwidth]{../figures/fig_repair_kscaling.pdf} +\caption{Menu-size scaling behind main-text Figure~\ref{fig:repair-acc}: the $K{=}4$ +checkpoint evaluated zero-shot at larger menus retains an accuracy edge at $K{=}8$ that +closes by $K{=}16$, while the blind enumeration it displaces grows from $2.5$ to $9.1$ +solver calls per instance.} \label{fig:repair} \end{figure}