Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

At this point, picking AMD for your CPU becomes such a no-brainer. Compounded by Intel’s security issues and all.


ECC support in AMD systems is strange. It's supported theoretically, but practically there are issues, one have to carefully pick motherboard and even then it's some kind of unsupported configuration. Intel sells cheap and fast Xeons with proper ECC support. I'm very interested in AMD CPUs and I hope that ECC story will improve, so I can buy some kind of workstation-branded motherboard and use fully supported ECC configuration.


The difference is that AMD doesn't disable ECC support in any model line, while Intel disables it, sometimes without rhyme.

Extra funny when you notice that certain Xeon lines are actually i7 with different branding and ECC left enabled.

The problems with ECC on AMD comes from consumer vendors not putting the time into testing, and possibly not even connecting the ECC lines (remember, ECC requires putting additional traces between memory controller and memory slots). Then you have to deal with whatever customisation the vendor of the motherboard did to firmware - their changes might have resulted in effective disabling of ECC.

With Intel, you either have the same game as above (with the non-Xeon ECC-capable parts), or pay through the nose for comparable performance "workstation/enterprise" gear, as ECC support being used for market segmentation by intel is pretty much an open secret.


You don't need to pay through the nose with Intel, at least for latest generation. 10900K costs $499. 10900K with ECC called Xeon W-1290P and costs $539. That's 8% extra. ASUS Pro WS W480-Ace is $280 which is reasonable cost for a good motherboard.


I picked up a i3-9100 a few months back because its was a low cost processor with ECC (for an edge/embedded solution). The problem then becomes the motherboard, and it seems intel has just shifted the ECC tax from the processor to the motherboard/chipset. That core fits on a lot of low cost motherboards, but to enable ECC requires about another $100 chipset tax.


How do you discover which Xeon is the i7 version of the same chip or vice versa? Is it just spelunking through Intel Ark to find a same generation, clock speeds Xeon or is there more to it?


Any of the Xeon Ws are just rebranded i7s, others might be as well. You can tell because they have the DMI memory bus instead of the Xeon exclusive UPI interconnect.


Something like that, yes. Specs are almost identical. There are CPU families so I picked best models from a family and compared them.


Would be interesting to compare die shots.


I was with you until the mobo cost. Wtf, $280 for a mobo?? A totally decent B550 mobo with ECC support can be had for less than half that price.

The only advantage of Intel Xeon at this point is RDIMM/LR-DIMM support on their Xeon-W platform. AMD supports that only on Epyc and Threadripper Pro.


Yeah, with Haswell you could get an i7-4770 for like $320 or a Xeon E3-1245v3 (includes an iGPU) that ran the same clock speed and spec, generally, but was about $270. If you went for the 1241v3 you lost the iGPU and 4 watts of power usage but gain +100mhz base clock.

And that gen of CPUS had low power T-series chips rated at 65w where today it's 35w. The power drop is part of why I'm happy with my SFF setup as the thermals are quite good on 14nm locked CPUs like my 7700.


ECC support is iffy at the consumer brands. Its a "we won't disable it, but we won't guarantee that it works" sort of deal.

If you want verified ECC support, you need to buy the workstation chips and motherboards: Threadripper Pro or EPYCs.


ECC is supported on the Pro series as well. My home server is running Ryzen 5 Pro 4650G (yay for integrated graphics) and Asrock B550M.

I went through the effort of using qvl memory, but actually testing ECC is a bit more difficult. While ECC is supported & active, memory errors are sadly not reported to the OS. I remember seeing a forum post somewhere of somebody overclocking/undervolting the ram to force errors, but I can't seem to find it right now. There's a fine line between stable, stable with recovered errors, and unstable.


That's what I'm talking about and I wouldn't call it "fully supported". I want to know about ECC statistics. It's important because if I can see that ECC recovers abnormally high number of errors, it's likely that I need to replace RAM right now.


ECC statistics are available for Ryzen if you use Linux (configured to load the appropriate EDAC module).

So on Linux, ECC is fully supported, even with Ryzen.


Do you know for which motherboards this is true? I think ASRock Rack X470D4U should be fine, as they market it as server board.

https://www.servethehome.com/asrock-rack-x470d4u-review-amd-...


I have used an older ASRock MB with the first generation Zen, and it was OK with ECC.

With Ryzen 3xxx, I have used ASUS Pro WS X570-ACE ($315), which is sold as a workstation board, so you definitely should expect ECC to work without problems, and also the Mini-ITX MB ASRock X570 Phantom Gaming-ITX/TB3 ($230), which also worked OK.

I expect that the other ASRock MB's (most of them or maybe all of them specify the support of ECC) also work OK with ECC modules.

The ASrock Rack server board should also not have any problems with ECC. IIRC that server board supports only up to DDR4-2933, but until now faster unbuffered ECC memory modules were not sold anywhere, so that is not a disadvantage.

Edit: I have looked again at the ASRock site and they have updated the memory support specification. If you use only 2 UDIMM modules (i.e. up to 64 GB total), then you can use up to DDR4-3200 (which I have never seen offered anywhere until now; 2666 is easy to find, 2933 is also supported by Intel since March, so it should become available soon).


Sadly, the IPMI implementation in the X470D4U series is awful. The remote console crashes frequently and is generally pretty unreliable. I'm disappointed that Supermicro doesn't have any Ryzen AM4 server boards. AMD getting ignored by many of the server vendors is just a repeat of what happened when the Opteron series was first introduced. At least there are a number of Epyc server boards available.


I don't disagree with you, but when was the last time you ever had to replace a stick of ram?


This year for me.


What operating system? On Linux, I had a memory stick that was not completely inserted, and periodically I saw corrected memory errors reported in the logs until I fixed the issue.


Indeed, forgot to mention this is on Freebsd.

https://forums.freebsd.org/threads/how-to-find-out-if-ecc-is...


The latest BIOS for ASUS Prime X370 Pro has ECC explicitly as a configuration option. Seems to work in Linux. I am using 2x8R ECC 2666Mhz RAM from Kingston.


Interesting! Do you think this BIOS or ECC works also on X370 Crosshair VI Hero? Hard to find info on this...


It's up to the motherboard vendor. They'll guarantee that it works if you use memory from the memory QVL.

Example: https://www.asus.com/no/Motherboards/Pro-WS-X570-ACE/HelpDes...


What Xeon and Xeon motherboard with ECC support are "cheap?"

In that price range, AMD markets Threadripper and Epyc, both with proper ECC support.

ECC support in Ryzen systems is up to the motherboard manufacturer, and some manufacturers advertise support very clearly. E.g., at least a couple years ago, ASRock explicitly supported ECC in all their Ryzen motherboards.


New workstation Xeons (Comet Lake), e.g. Xeon W-1250, W-1270, W-1290 are hard to find online but can be seen, in few places, for 350,500,800 euro. Motherboards for these Xeons (socket LGA1200) seem to be a somewhat pricier at 300 euro, but they positively have ECC. Intel still costs more for same perf/capabilities, but if you need ECC and Intel platform, it isn't crazy.

Most people wanting ECC just for the warm fuzzy enthusiast feeling would prefer Ryzen with good MB that supports ECC. Intel makes more sense at the high-end many-core Xeons, due to supporting huge memory with much higher throughput (LGA3647). But that segment is bonkers-expensive.


That's an interesting kind of enthusiasm. The usual PC enthusiast thing is to overclock memory to hell until the system starts glitching and dial it back just a little and call it a day :)


I don't feel myself confident with flaky hardware. I'm using UPS and ECC, so I can be sure that no bugs are introduced by hardware issues. I would love to use GPU with ECC, but they are ridiculously priced, so I'm out of luck.

That's not a real need, of course, it's just some kind of whim, but if I can pay for it, why not. Just like those crazy overclocker dudes pay for their whims.


You are right, but if you buy a motherboard that claims to support ECC, you will usually not have any problems.

For example I am using an ASUS Pro WS X570-ACE, which is a reasonably priced workstation board ($300) with a Ryzen 7 3700X and ECC memory.

ECC worked OK, without any problems. I have also used a couple of ASRock MB's and ECC also worked OK on them.

I would much prefer more guarantees from AMD, but rather than buying a slow Intel CPU I prefer a little risk with AMD.


AMD guarantees ECC works with their workstation focused Threadripper line. On Ryzen, it only works if you do your homework picking hardware.

It's a shame they made the TR platform much more expensive in the last generation.


Is that a new thing? I have an older Threadripper and my motherboard definitely doesn’t support ECC (even though the CPU does)


As far as I know all Threadripper chips and board were required to support ECC. Mind you this means UDIMM ECC commonly targeted at workstations, not the RDIMM ECC used in servers.


Dumb question. Why would one want ECC in something that isn't a server? How often do bits in memory actually flip by themselves for it to be warranted?


According to Google statistics, bit flips are encountered at a rate of 8% per DIMM per year. https://research.google/pubs/pub35162/

You want ECC in your desktop if you cannot afford your data to become (silently) corrupted. A single bit-flip in a JPEG image can totally ruin it. And you won't notice without opening it, because the thumbnail is unaffected. When you finally notice, the corruption has possibly spread to backups already.

For servers that is especially important, because data will reside cached in memory for days or even weeks. Also ECC is often the last line of defense against Rowhammer.


https://www.cs.virginia.edu/~gurumurthi/papers/asplos15.pdf is the latest study I can find. Some key findings; error rates per DRAM module and per technology cycle (DDR2, DDR3, and so presumably DDR4) have similar failure rates. Higher altitude has higher failure rates.

FIT (failures in time, or failures per billion hours) are about 20-30 per DRAM (18 per DIMM in the study) at a fault level, which may not rise to an error condition. ECC corrects many faults, if present, which don't become errors.

So expect a (4GB in this study) DIMM to have a fault once per 200 years on average. Of course, this is server-class hardware running ideal conditions.

This is, notably, in stark contrast to Google's results (https://research.google.com/pubs/archive/35162.pdf&) of about 20,000-70,000 FIT per Mbit, but closer reading suggests that there were a fairly large number of high-fault DIMMs affecting the mean and Google historically bought the cheapest hardware available. Even at this fault rate an average DIMM wouldn't fail in any given month, for example.

Personal gaming, browsing and light document work PCs do not need ECC. Video/photo editing/development PCs might be on the threshold for not wanting to produce or store corrupted work.

Multiply the risk out for any organization with many people doing the same kind of work where corruption on a single workstation could cause annoying trouble for everyone else.


Good question. Most people don't really need it, as systems work quite well for long enough time without it and no visible corruption happens. But if you are tech enthusiast who knows a thing or two about errors then you want it, because it makes your system "serious"-grade. It boosts your confidence in quality of your system, and thus self-confidence, ego and bragging rights. Also, it makes your system more professional-grade, where you can monitor ECC errors and, if needed, replace RAM modules or motherboard.


https://stackoverflow.com/questions/4109218/do-gamma-rays-fr...

tl;dr: one bit error in 4GB every 72 hours


I wish someone with a larger server farm would count the number of reported ECC errors per GB-hour and give us updated numbers. That StackOverflow question is about 10 years old now, and I think it's relying on data even older than that.


Someone once did a bit-squatting experiment and "estimates that 614,400 memory errors occur per hour globally".

https://nakedsecurity.sophos.com/2011/08/10/bh-2011-bit-squa...

It would be interesting to repeat this experiment today.


Yes; per the internet archive[1], the data's at least 20 years old.

1. https://web.archive.org/web/20010612184424/http://www.boeing...


A very good way to undercut low end Xeons which were bought solely for their insurance against an out of the blue crash.


You do have to pick a motherboard that supports it, but when you do, it should work without issues as long as your CPU supports it. Not all ryzen CPUs support ECC.

Here is an example: ASUS Pro WS X570-ACE


Aside from that one example, are there others? I'm not a fan of its fan.


I'm searching now but does AMD have an alternative/answer for Intel's QuickSync? Turning on HW acceleration on my Plex server (so that it uses QuickSync) is a game changer. From struggling to handle 3+ 1080p streams and pegging all the cores to being able to do 6+ without going over a load average of 1.


Video encoders/decoders are parts of GPUs, not CPUs. Only CPUs with integrated GPU have these pieces of hardware.

AMD is pretty comparable in that regard: https://en.wikipedia.org/wiki/Video_Core_Next but I don’t have computers with AMD APUs or GPUs, and don’t have a hands-on experience with these features.


Hmm, looks like Plex doesn't have support for Video Core Next yet: https://forums.plex.tv/t/feature-request-add-support-for-amd...

I'm still probably 3-6mo away from a new server build so I'll just re-evaluate then I guess and honestly I might just go with another storage server and leave my intel/QS server as-is and just go with a AMD that plays nice with UnRaid.


Which OS are you running there? If it’s Windows, does your software have an option to select MS Media Foundation for video encoding/decoding?

In my experience, GPU vendors, all 3 of them, are including reasonably well-made media foundation hardware transforms as a part of their GPU drivers. MF API is vendor agnostic. Apart from a few bugs I found in Intel’s drivers (it was about h265 hardware encoder, Intel neglected to react), the same API works with all capable hardware.


UnRaid (syslinux under the hood IIRC) and I'm running Plex in Docker. I've seen guides on doing a GPU passthrough to the docker container so I know it's possible to do. That might be my next build, a AMD CPU and a GPU for Plex, it will depend on where UnRaid AMD support is at the time (it has had issues in the past from what I've seen on the forums) but I really want my next build to use AMD.


Plex works so well with Ubuntu that I have become a huge proponent of this method. I’m not sure if it’s their developers having a bias for the OS, but the Ubuntu version always seems to work well.


Plex supports Nvidia GPUs for hardware transcoding, so you could pick up a cheaper Nvidia GPU and stick it in the build, but that will probably not be possible for small builds.


I wanted AMD but quicksync for a Plex server is just so good. I bought an I7 NUC10 for this role and it’s great. Virtualised OS, Docker for Plex and it’ll do 11x transcodes (1080p to 720p) in hardware while also hosting several other machines. The first time you pass though the GPU to the Vm, then into Docker is a bit of a head scratcher, but it’s actually fine and works well.

It has a maximum of 64gb ram and is tiny. With an nvme drive (Samsung Evo) it is really really fast.

The next best option as far as I was concerned was the NUC8, and while I’d love to have the PSU onboard and no brick, a Mac Mini is a lot of money.

The Nuc8 is a better option than the 10 for anything that’s needing actual graphics. The 10 has a very anaemic GPU compared to the 8, but has 2 more CPU cores when comparing i7s.


I hadn't considered a NUC to run the plex server, right now I enjoy managing everything through UnRaid but I'm not opposed to another device that would manage plex alone. QuickSync is CRAZY good, I stumbled onto a post about it a month or so ago and thought "Ehh, I'm probably already using it, w/e" then saw I needed to run some commands and mount a device in my plex docker container to actually enable it and I was blown away by how little load there was on the CPU. I know there has been some grumbling about QS not producing the best output but that seems to be mainly related to the earlier iterations of it, I have not noticed any issues. It kills me how long I waited to try this, it literally took <10min to setup and has been a huge performance boost.

To everyone else: If you have an Intel CPU that supports QS and you are running Plex in a docker container then I HIGHLY recommend you look into enabling HW acceleration (both in plex settings and by mounting a device, /dev/dri, into your container). Here [0] are the directions I followed. Make sure you read the instructions, the first time I tried to follow this I misread the instructions and thought my CPU didn't support it or there was some other issue but it couldn't have been easier when I actually stopped and read instead of skimming.

[0] https://forums.unraid.net/topic/77943-guide-plex-hardware-ac...


This is similar to my experience. There are some great guides on Reddit too. I found it a bit tricky in a VM as the ESXi default kept taking over and breaking it. I figured this out by poking about from within the container and discovered it had two GPUs and was picking the wrong one. Once I’d deleted the card VMware made (which means you can’t see the console in ESXi GUI interface, but can still ssh in) it all started working. As you say, it’ll do a lot of work, CPU usage is almost nil and I can’t see a quality issue.

https://www.google.com/amp/s/amp.reddit.com/r/PleX/comments/...

> I hadn't considered a NUC to run the plex server

Mine is more of a Docker host. Home Assistant absolutely flies on that hardware (it was previously on a Synology 918). It has a few other well used containers too - Wireguard, several Minecraft and some administrative tools. There are several testing machines for pfsense and some Linux machines I wanted to play with too.


QuickSync is GPU accelerated encode/decode right? This processor announcement is for their CPUs without GPUs, so you'd need a GPU add on board, and both AMD and NVidia support that. AMDs processors with GPUs (they call them APUs) support that too. AMD tends to release desktop CPU, then high end desktop/server, then laptop APU and finally desktop APU. They only released Zen2 desktop APUs a couple months ago, and they're currently OEM only and very hard to find in the US (grey market imports only, AFAIK, but send me an email if I'm wrong, address in profile)


QuickSync is all in the CPU (no GPU needed). IIRC it's part of the Intel Graphics (built into the CPU) so maybe it's not exactly fair to call QuickSync part of the CPU but it's included in the physical CPU chip and I have no discrete graphics card in my server right now. I know I can get a GPU to offload decoding/encoding to but QuickSync is pretty awesome for my use-case and buying a graphics card has it's own issues (space in case, cost of card, getting it play nice with Plex in docker, etc).


It is GPU needed. The intel spec sheet for i3-9350KF doesn't show QuickSync, but i3-9350K does. The difference between the two is that the F series doesn't have a GPU (or it's disabled). Also unavailable on server class Xeons without GPUs.

I agree, it's convenient to have a GPU on most CPUs, but AMD puts less priority on that market, so we just have to wait.


When you have 24 "HT" cores I'm not sure why you would need QuickSync.


Because in addition to theoretical capacities I also care about my power bill and how comfortably cool the room my transcoding machine is in stays.


Because video resolution has grown well above 1080p. I’m looking at 4k monitor at the moment, recent chips have some support for 8k video.


For a normal machine that does look to be the case but I've always found AMDs manuals and software quite lacking so it may be worth going with intel just for the tooling (i.e. performance counters seem to be much better documented on intel)


It's not that simple for computing. I heard that in Data Science Intel is still preferred because of better AVX support.

There are also things like Intel MKL. A lot of software can use it when compiled on a user machine.


> is still preferred because of better AVX support

AVX1 and AVX2 performance is on par.

For instance, vmulpd AVX1 instruction is faster on AMD, 3 versus 4 cycles. vpaddd AVX2 instruction is same at 1 cycle latency. vfmadd132pd FMA instruction is slightly faster on Intel, 4 versus 5 cycles. Throughput is the same across these two. I was looking at AMD Zen2 versus Intel Ice Lake.

Some Intel chips have AVX512. Still, many practical applications don’t need that amount of SIMD wideness, and these who do are often a good fit for GPGPUs.

> There are also things like Intel MKL

There’re vendor-agnostic equivalents like Eigen.


IIUC, Intel uses the term "AVX-512" as an umbrella term, and different processors support different subsets of "AVX-512" instructions [0].

AFAIK this is a break from previous Intel nomenclature, where any processor supporting e.g. "SSE4.2" instructions was guaranteed to support all SSE4.2 instructions.

I'm concerned that sometimes this causes confusion when talking about processor -- software compatibility.

[0] https://en.wikipedia.org/wiki/AVX-512


I think the consideration GP was trying to make is that Zen (at least 1 and 2 — and I haven't heard otherwise for 3) do not support 512-bit wide AVX registers at all.


> There’re vendor-agnostic equivalents like Eigen.

That's looking at the wrong layer of the hierarchy, I think. There are many open-source linear algebra libraries, but iirc they all link against something that has a BLAS/LAPACK API. That might be something like MKL, OpenBLAS, ATLAS, etc.

When I last checked, MKL was much faster than its competitors, and is only available (at full speed) on Intel CPUs. Has that changed?


> iirc they all link against something that has a BLAS/LAPACK API

Eigen can consume these I think, but they are optional. It has it’s own implementation of these, written in manually vectorized C++, with intrinsics, up to and including AVX512 (controlled with macros). For parallelization it uses OpenMP provided by the compiler (also controlled with a macro).

> Has that changed?

It’s hard to directly compare Eigen to the rest of them. They don’t do the same thing.

One feature of Eigen is lazy evaluation. Expressions like a+b or a×b don’t return another matrix or vector; they return a placeholder object that only computes something on assignment. For complicated expressions this can be a huge win, e.g. r=a+b+c+d will read from a,b,c,d, compute sum of the 4 on the fly, and write into r without temporary copies in memory.

However, also makes Eigen’s source code outright scary, and hard to debug or optimize.

Anyway, based on the old pics there http://eigen.tuxfamily.org/index.php?title=Benchmark they are more or less comparable. Things like alpha·X+beta·Y were much faster in Eigen (probably due to that lazy evaluation thing), Hessenberg was much faster in MKL, in general they are close.


> When I last checked, MKL was much faster than its competitors

That has never been generally true in my experience measuring over the years. It has been true at times for specific cases, e.g. OpenBLAS until it got avx512 support on a par with MKL (at least for serial DGEMM -- I've forgotten quite how the rest of level 3 goes).


For some computer vision tasks, Tensorflow is much faster when you have AVX512.

Also https://www.intel.com/content/www/us/en/artificial-intellige...


These results were achieved on dual-socket Xeon E5 2699v4 (the architecture is 5 years old and has no AVX512, they optimized for AVX2) and on Xeon Phi 7250 (that thing does have AVX512 but that’s not a processor, a specialized accelerator with 68 cores).

Also Tensorflow is awesome fit for GPGPUs and is usually way faster on them.


The new Zen 3 cores are expected to have a higher AVX throughput per cycle than all Intel CPUs, except the most expensive models of Xeon Gold, Platinum or W and the HEDT i9 models that have dual AVX-512 FMA units.

The cheaper models with only one AVX-512 FMA unit have a lower throughput, which will be exceeded by Zen 3, even at the same clock frequency.

For multi-threaded tasks, Zen 3 CPUs will have a higher clock-frequency than any Intel CPU, so it is expected that any older Intel CPU will be beaten easily.

It remains to be seen which will be the performances of the Ice Lake Server CPUs, to be launched before the end of the year. However, miracles are not expected, because these are using the older Intel 10 nm technology, not the improved one used by Tiger Lake.


> The cheaper models with only one AVX-512 FMA unit have a lower throughput, which will be exceeded by Zen 3, even at the same clock frequency.

I think there are relatively few with only one FMA (and there's no way of interrogating them at runtime, sigh) but, yes, if you know you have one, you use AVX2 for GEMM kernels, as a specific example.

For general computational workloads, you're likely better off with more AVX2 cores and high memory bandwidth, even without whatever improvements there are in Zen 3.


I've ordered a tiger lake laptop which should arrive around the end of the month, so I'll be able to test then. So just speculation for now:

I'd think AVX512 would still be advantageous for gemm kernels. If you use a 16 x 14 microkernel instead of an 8 x 6 microkernel, you'll reduce memory bandwidth by halving the passes over the cache-sized blocks in memory. Well optimized implementations don't really have a problem with this (getting close to the CPU's peak flops on AVX2), but it still seems better than not doing it, and is an advantage for code with register tiling more generally.

I suspect that it'd help people using Tullio.jl in Julia, for example.


Yes, even on a computer where the throughput of AVX-512 is the same as the throughput of AVX/AVX2, like Tiger Lake, it is usually much easier to reach that throughput with AVX-512 than with AVX/AVX2.

The mask registers can eliminate special prologue or epilogue code for the loops and having more and larger registers make it much easier to overlap enough computations so that the latencies of the operations are hidden.

I would never buy again an Intel CPU without AVX-512, because that has already for some time been their main advantage and now it remains their only advantage over AMD.

Unfortunately for Intel, even if AVX-512 is a great improvement, it cannot compensate for a number of cores half of the competition. Intel needs to launch the 8-core Tiger Lake H no later than March 2021, but it would be better for them if they could do it earlier.


> it is usually much easier to reach that throughput with AVX-512 than with AVX/AVX2

That's not what I understood from people who've gone through the exercise for GEMM. One probably relevant measurement: on BLIS' generic C GEMM micro-kernel, GCC doesn't get nearly as close to the hand-coded avx512 version as for avx2 with appropriate bock sizes.


Yes, if you have two AVX512 FMA units, you use them for GEMM; with a current microarchitecture implementation, if you only have one, you don't. https://github.com/flame/blis/blob/2d8ec164e7ae4f0c461c27309... GEMM is one of relatively few computations with sufficient computational intensity.


Most or all workstation CPUs with AVX-512 (Xeon W or Core i9) have indeed dual 512-bit FMA units.

However most server models with AVX-512 have only one 512-bit FMA unit. That includes all Xeon Bronze, all Xeon Silver and a few Xeon Gold.

All Xeon Platinum and most Xeon Gold have dual units, but those are extremely expensive.


Fair enough. I've just never come across 1 FMA versions, despite trying to enumerate them -- doubtless computational bias.


> and there's no way of interrogating them at runtime, sigh

Well, directly, no.

Indirectly ... find the model and use a lookup table to find how many FMA units it has :P


I've been through the exercise. It's obviously tedious and fragile, and not always possible, as far as I remember. I can't give an example of something that wasn't listed, but I at least have queries against D series and i7/i9.


> There are also things like Intel MKL

There are also things like OpenBLAS, and BLIS (which AMD support).


Unless you want an integrated video card in your high-end CPU.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: