My first impression is that it's not so good at following prompt directions. I asked it to place a 3D text made of glass in a particular city. It instead gave me a broken 3D text on a white background. Maybe with different seeds it gets better, but it's more of a trial and error process than reliable results.
You could try attaching other images as references (I think you can attach a maximum of 10 images). If the attachments can be blurred or sketchy or generic enough, they could be used for generalization.
docheinestages · · focus · HN ↗
dannyw · · focus · HN ↗
mdp2021 · · focus · HN ↗