Both GPUs must be dedicated to the LLM for 262K context window. Image gen and LLM run one at a time (swap takes seconds). Simultaneous only possible with smaller context (~64K on one card). https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu
Both GPUs must be dedicated to the LLM for 262K context window. Image gen and LLM run one at a time (swap takes seconds). Simultaneous only possible with smaller context (~64K on one card). https://claude.ai/code/session_01PtYTPherSJaxDEVPgF6Nxu