馃 AI text below 馃
Observation
The SRC sweep sketches with exactly chi_out columns and returns that sketch as the output bond. A spike for #34 showed that this accounts for most of the error, and it dominates any gain from compressing a whole stack in one shot. Relative errors against the exact product, compared with a lower bound (the largest best-rank-chi tail over all cuts, which no chi train can beat). Setup: 10 sites, 20 seeds, median values:
| case |
chi |
lower bound |
SRC, sketch chi |
sketch 2 chi + exact truncation to chi |
| random MPO^3 . MPS |
8 |
9.8e-2 |
2.7e-1 |
1.4e-1 |
| random MPO^3 |
16 |
1.0e-1 |
3.8e-1 |
1.8e-1 |
| Trotter U^3 . MPS, dt=0.3 |
4 |
2.1e-3 |
9.2e-3 |
2.6e-3 |
| Trotter U^3, dt=0.3 |
8 |
1.7e-7 |
2.4e-6 |
2.0e-7 |
| Trotter U^3, dt=0.8 |
8 |
3.2e-4 |
3.5e-3 |
3.6e-4 |
Without oversampling, SRC sits 3-10x above the bound. Sketching at 2 chi and then truncating exactly to chi brings it to within 1-1.5x of the bound in every case measured. The same holds for sequential pairwise apply (each step oversampled and then truncated).
Proposal
Add an oversample factor (or an additive p) to apply and compress:
- Run the existing sweep with sketch size
ceil(oversample * chi_out).
- Truncate the result to
chi_out with one left-to-right SVD sweep.
The SRC output is already right-canonical: every site except the first is a right isometry built from the orthonormal columns of Q. So the SVD of each centre tensor gives the exact Schmidt values at that cut, and no re-orthogonalization is needed:
# padded sites (l, r, u, d); train is right-canonical, norm on site 0
for j in range(n - 1):
l, r, du, dd = sites[j].shape
U, S, Vh = svd(sites[j].transpose(0, 2, 3, 1).reshape(l * du * dd, r))
k = min(chi_out, S.size)
sites[j] = U[:, :k].reshape(l, du, dd, k).transpose(0, 3, 1, 2)
sites[j + 1] = contract("ab,bcud->acud", S[:k, None] * Vh[:k], sites[j + 1])
Each cut is truncated optimally for the state as truncated so far, and the discarded weights add in quadrature, which gives the usual TT-SVD quasi-optimality. The pass costs O(n p^2 chi_os^2 chi), which is small next to the sweep.
Open questions
- Output canonical form. The pass above returns a left-canonical train, but
apply and compress currently return right-canonical ones. Either document the change or add a cheap QR sweep to restore right-canonical form.
- Interaction with
cutoff. Today cutoff truncates on the singular values of the sketch's R factor. With the exact pass, a relative cutoff on true Schmidt values becomes available and is arguably what users expect. Should cutoff move to the final pass when oversampling is on?
- Cost of the larger sketch. The spike measured accuracy only. The sweep cost grows at least quadratically with the sketch size, so we need an accuracy-per-second comparison against simply raising
chi_out.
- Default value. Keep
1.0 (current behaviour, bit-for-bit) or pick something like 1.5-2.
- GPU. The pass must go through
xp like the rest of the sweep.
The spike code (throwaway, not in the repository) compared against dense references on 10-site chains; see the discussion on #34.
馃 AI text below 馃
Observation
The SRC sweep sketches with exactly
chi_outcolumns and returns that sketch as the output bond. A spike for #34 showed that this accounts for most of the error, and it dominates any gain from compressing a whole stack in one shot. Relative errors against the exact product, compared with a lower bound (the largest best-rank-chitail over all cuts, which nochitrain can beat). Setup: 10 sites, 20 seeds, median values:chichi2 chi+ exact truncation tochiWithout oversampling, SRC sits 3-10x above the bound. Sketching at
2 chiand then truncating exactly tochibrings it to within 1-1.5x of the bound in every case measured. The same holds for sequential pairwiseapply(each step oversampled and then truncated).Proposal
Add an
oversamplefactor (or an additivep) toapplyandcompress:ceil(oversample * chi_out).chi_outwith one left-to-right SVD sweep.The SRC output is already right-canonical: every site except the first is a right isometry built from the orthonormal columns of
Q. So the SVD of each centre tensor gives the exact Schmidt values at that cut, and no re-orthogonalization is needed:Each cut is truncated optimally for the state as truncated so far, and the discarded weights add in quadrature, which gives the usual TT-SVD quasi-optimality. The pass costs
O(n p^2 chi_os^2 chi), which is small next to the sweep.Open questions
applyandcompresscurrently return right-canonical ones. Either document the change or add a cheap QR sweep to restore right-canonical form.cutoff. Todaycutofftruncates on the singular values of the sketch'sRfactor. With the exact pass, a relative cutoff on true Schmidt values becomes available and is arguably what users expect. Shouldcutoffmove to the final pass when oversampling is on?chi_out.1.0(current behaviour, bit-for-bit) or pick something like1.5-2.xplike the rest of the sweep.The spike code (throwaway, not in the repository) compared against dense references on 10-site chains; see the discussion on #34.