Rather than individual-memory-address prefetch instructions, how would you feel about sending a DMA program to a controller on-board a memory DIMM, that would then enable you to send short external commands to the memory that would be translated by the DMA program into custom “vector requests”, to which the memory could respond with long streams of fetch responses—shaped somewhat like the output of a CCD’s shift-register—where this stream of fetch responses would then entirely overwrite the calling CPU’s cache lines with the retrieved values?