Well, in general the file (png/jpg) is loaded into memory an uncompressed and finally copied into video memory as a texture. In the browser/construct world the image file is loaded into an html img element before it's copied over to a webgl/webgpu texture. Presumably the browser may also load the image to a different texture too if it were to display the image onto the page outside of construct's renderer. However, the exact details of the browser renderer and even construct's renderer are a black box.
But we can still calculate how much video memory an image would take up. For example, if a 4k image is 4096x2160 then the uncompressed size would be approximately 34mb. (width*height*4/1024/1024).
1. Seems reasonable that loading three images instead of one would use more memory. For the html image to display it has to load into vram before the browser can render it. The question as to why the amount of vram usage is bigger than we calculate is largely due to the black box nature of the browser and construct's renderer. Maybe they need more textures to composite stuff or whatever.
2. Yes, when you clone the object types it duplicates the images. Guess you could make a request to have construct check to see if images are duplicated to save memory but that would slow down export times a lot.
3. Yes, but you could have tested that. Instances of the same type will share textures.
4. Presumably it loads a new texture.
5. It's likely more performant to just render everything onto the canvas instead of layering the canvas with html images. Also, you can't do the same things with html images as you can do with sprites. But luckily, you can do whatever you like.