Source
engines/quan_loncar_meep_lab/notes/device23_seed_lineage.md · assembled 2026-09-05 15:22 UTC.
Device23 launch seed: preserve the device before optimizing it¶
The inherited design23_20260801/undercut_mirror.json is useful as a 1550 nm
layer-stack control, but it is not Device23. It is a 900 x 220 nm undercut
beam with eleven identical holes, including one at x = 0; it has no defect
and its measured localization was 0.21.
The supplied Device23 fabrication notebook instead defines an oxide-backed,
795.692787 nm cavity in a 556.263652 x 300 nm SiN beam. It has 42 positive
holes reflected across x = 0 (84 total), no central hole, a
234.625430 nm pitch, 22 tapered pairs, and 20 mirror pairs. The recovered
FDTD result is Q = 11,812.391. That exact lineage control is exported as
benchmarks/design23_20260801/device23_canonical_84hole.json.
The full article is not the practical local launch seed. A later Device23 campaign retained the inner 20 pairs and physically evaluated its accepted step 6 in Tidy3D. The checked result is:
| observable | qualified value |
|---|---|
| wavelength | 770.029669 nm |
| Q | 1,113.003 |
| anthracene-normalized V | 6.180184 (lambda/n)^3 |
| Q/V | 180.092 |
| two-port TE0 beta, power | 0.878603 |
| two-port TE0 beta, Q cross-check | 0.886978 |
| TE0 fraction of x-directed flux | 0.962033 |
| requested/reported lead modes | 4 (not exhaustive) |
All twelve coherent-observable checks pass, including incoming-wave rejection,
power closure, on-resonance alignment, TE polarization, and the independent
power/Q beta cross-check. The coherent Tidy3D simulation requested four
modes, so its four reported effective indices are provenance, not proof that
the lead has exactly four guided modes; local qualification must search for
additional weak branches. Its exact 62-coordinate geometry is exported as
benchmarks/design23_20260801/device23_40hole_fdtd_seed.json; this is the
free-form campaign seed. The complete source is
engines/design23_v1/runs/20260728T013605Z_design23_40hole_composite_bfgs/
fdtd_step_0006_diagnostic/diagnostic_result.json.
Stack translation, not a stack change¶
The legacy coordinates are translated upward by 150 nm for the new
LayerStack convention:
| material | legacy z (um) | layered-engine z (um) |
|---|---|---|
| SiN | [-0.15, 0.15] | [0.0, 0.3] |
| anthracene | [0.15, 0.35] | [0.3, 0.5] |
| PVA | [0.35, 0.55] | [0.5, 0.7] |
| anthracene emitter | 0.25 | 0.4 |
The oxide substrate is retained. It is physically semi-infinite; the JSON
bottom_um = -0.6 is only a numerical truncation, and the stack builder extends
the bottommost unetched substrate to the lower physical/PML boundary. The
launch seed is deliberately not undercut because no undercut version of
this geometry has equivalent pole and modal evidence.
Exact 770 nm native lead certificate¶
The substrate-backed step-6 lead was qualified locally at its exact 770.029669 nm target, independently of the unrelated 1550 nm undercut control. The mode solver uses a full-y/full-z vector Yee Bloch cross-section and a one-plane Lorentz-reciprocity E/H projection. A 25/um solve with a 20/um coarse control, 24 beta samples, 30 eigenpairs, and deep port padding of 0.6 um in y and 1.0 um in z produced the following evidence:
| control | qualified value |
|---|---|
| stable localized branch count | 6 |
| selected TE0 effective index | 1.778786891 |
| selected TE0 Ey fraction | 0.909672346 |
| target beta grid difference | 0.019189852 |
| target boundary energy | 0.000051042 |
| target reciprocity self error | 4.44e-16 |
| opposite-direction leakage | 4.30e-17 |
| maximum other-branch overlap | 2.57e-10 |
The selected TE0 is branch 1 rather than the largest-beta branch because branch 0 is TM-like. The historical Tidy3D monitor requested exactly four modes, so the six local branches do not contradict that capped report.
The deeper two-dimensional mode box is deliberately larger than the pilot's
three-dimensional physical cross-section. Before sampling the 3-D field, the
TE0 functional is cropped to y = +/-0.778131826 um and
z = [-1.05, 1.15] um. Its retained self-overlap is 0.999913692; the omitted
8.63078e-5 is below the pinned 1e-3 gate, after which the retained
functional is renormalized to unit target overlap. Sampling therefore never
enters the transverse PML.
Weak unwanted branches close to the substrate light line have 7--9% boundary
energy even in the deep box. This is a recorded nonfatal warning, not a
numerator qualification failure: the hard numerator projects only the
well-converged TE0, and the physical full-y dipole-work denominator counts all
other guided and radiative power as unwanted. Fine/coarse branch-count
stability remains mandatory so a missing branch cannot be hidden by that
policy. The complete 20/16 failed control and passing 25/20 follow-up are in
benchmarks/design23_20260801/device23_770_layered_port_control.json.
Geometry and launch gates¶
On the 4 nm level-set grid the raw seed has 40 voids and a 77.695 nm minimum feature. The production 12 nm wall filter, one-cell step filter, reinitialization, and 80 nm island-area pruning retain all 40 voids; the settled field has a 76.878 nm minimum feature, 109.810 nm longitudinal gap, and 146.364 nm sidewall web. Therefore the free-form feature floor is 70 nm. The 80 nm island setting is safe because it is an area-equivalent threshold, not a minor-diameter threshold.
The default adaptive 96-plus-node spline is too expensive on forty level-set
voids. A fixed-node audit against 2,048 samples per void is recorded in
benchmarks/design23_20260801/device23_aperture_resampling.json. At 32 nodes,
the conservative pre-spline polyline has a 1.790 nm maximum Hausdorff error and
0.867% maximum area error. At 48 nodes those fall to 0.987 nm and 0.412%, while
the 76.878 nm minimum feature and 109.810 nm gap are unchanged. The named
Device23 pilot fidelity therefore uses 48 fixed nodes per void and disables
only the adaptive node multiplier; the global 96-node default is unchanged.
This aperture approximation still needs a 32-versus-48 pole comparison before
it can become a certification fidelity.
The resulting fixed-48, all-order-one mesh has 619,474 degrees of freedom and
builds its mesh/space in 143.56 seconds under concurrent eigensolver load; the
machine-readable record is device23_pilot_mesh.json. In contrast, the named
medium mixed-order mesh reaches 2,072,871 unknowns even with exact OCC ellipses
and is refused by the 500k production safety gate before eigenanalysis. This
is why the campaign labels the tractable setting device23_pilot_p1 and does
not present it as an order-converged Q/V result.
The corresponding physical full-y point-source domain was also built without
assembly or factorization. At the same fixed-48, all-order-one fidelity it has
1,252,734 degrees of freedom, builds in 70.27 seconds, and uses 722,956 kB peak
RSS without swap. This expected near-doubling is recorded separately in
device23_pilot_full_y_mesh.json; it is a resource measurement, not an LDOS,
volume, or coupling result.
That pilot then failed the pole smoke. With the exact OCC ellipses (so aperture
resampling was not involved), the closest candidate was
1.2938 - 0.0085i /um, residual 4.3e-6 and legacy central-patterned-energy
selection value 0.176. None of eight candidates met the 1e-8 residual gate;
the largest legacy selection value was 0.252. Those values are not comparable
to the later Ey-monitor localization. The result is recorded in
device23_40hole_stack_pole_smoke.json. Consequently
device23_pilot_p1 is a mesh-scaling diagnostic only and is not safe to
launch as a Q/V or coupling-restoration optimizer fidelity.
Increasing the Arnoldi request from 8 to 16 candidates repaired convergence
but not the fidelity. On the settled fixed-48 field it returned residual
1.14e-13 at 765.439 nm and Q = 120.64. Its reported 0.2635 selection value is
the legacy central-patterned-energy fraction because the process imported the
solver before the Ey-monitor selector update; it must not be compared to the
current Ey localization metric. The run used
619,474 DOF, 342 seconds, and 17,562,636 kB peak RSS without swap. Since the
coherent FDTD seed has Q = 1,113, this converged Q = 120 branch still fails the
fidelity gate and cannot launch the optimizer. The full record is
device23_40hole_stack_pole_p1_c16.json.
A selective-order mesh ladder on the same fixed-48 settled field gives:
| half-y fidelity | DOF |
|---|---|
| base/PML order 1, cavity order 2 for abs(x) < 2.8 um | 1,447,138 |
| base/PML/cavity order 1, inner order 2 for abs(x) < 1.4 um | 972,976 |
| base order 1, PML order 2, cavity order 2 for abs(x) < 1.4 um | 1,004,130 |
The last case is the smallest ladder point that raises both the central cavity
and the radiation-sensitive PML while staying near one million unknowns. It is
the next scientifically defensible pole fidelity if a larger all-order-one
Arnoldi space still cannot recover the branch. The exact settings and mesh
times are stored in device23_order_ladder.json.
The seed begins below the requested 95% two-port TE0 constraint. It must enter the optimizer through the measured restoration phase, not be mislabeled as feasible. Launch remains fail-closed until the local layered pole and E/H feedthrough projections reproduce a localized branch and a physical modal budget. Raw x-PML absorption cannot satisfy that gate.
Channel resource envelope, 2026-08-02¶
The two-port feedthrough is measured on a full-y domain because a point current on a PEC mirror has no stable H(curl) normal trace. That domain is 1,252,734 unknowns against the pole's 619,474, and it is driven rather than an eigenproblem, so it factorizes where the pole iterates.
Measured, not estimated:
| solver | limit | outcome |
|---|---|---|
| UMFPACK | 38 GiB virtual | numeric factorization out of memory |
| MKL Pardiso | 36 GiB resident | OOM killed at 36 G peak, 15m54s wall, 40m CPU |
Pardiso is the better tool -- multithreaded, far less fill -- and it still does not fit in 36 GiB on a 45 GiB machine at the pilot channel fidelity.
Two practical notes for anyone repeating this. ulimit -v is the wrong
guard: it caps virtual address space, and MKL with its OpenMP arenas
reserves generously enough that Pardiso reports out-of-memory while
resident use is far below the cap. A cgroup limit through
systemd-run --user --scope -p MemoryMax=... -p MemorySwapMax=0 caps
resident memory instead, which is the quantity that actually threatens the
machine, and it kills only its own scope. And installing MKL reaches
further than the call sites that name a solver: ArnoldiSolver selects the
default sparse solver itself, so the pole solve moved to Pardiso without
being asked.
Python buffers stdout to a pipe, so a run killed mid-solve reports nothing
at all unless invoked with -u. The first attempt lost its completed pole
result that way.