Comment

Comment on LPT Do it.

CCF_100@sh.itjust.works ⁨1⁩ ⁨year⁩ ago

docx files are actually zip archives with xml in them

source

Sort:hotnew top

petersr@lemmy.world ⁨1⁩ ⁨year⁩ ago
Let me tell you something. I cannot tell you what company, but I have been tasked with putting Excel files in git “because they are just zip archives with xml” and it is just a disaster. Everytime you save the document it will save certain parts of the xml code in arbitrary ways (like each image is in a list and the order Og that list is random everytime), some metadata is re-written everytime like time of last modified and finally all the xml files are one single line. The git diffs are complete useless and noisy and just looking at the Excel file will cause git to consider it updated. So sure, you can use git to snapshot you Office documents… But just don’t.

source
- lengau@midwest.social ⁨1⁩ ⁨year⁩ ago
  If you are, like I once was, the poor fool who has to maintain a bunch of VBA macros… Extract them into files and source control those. Make a script to extract them and to put them back, and use git-lfs for the actual workbook if you need a template workbook.
  
  Now pardon me, I need to add this to the agenda for my next therapy.
  
  source
  - petersr@lemmy.world ⁨1⁩ ⁨year⁩ ago
    I will join that therapy session. This is pretty much what we did, except LFS, since it was “a requirement” to also track what they layouting of the Excel file was like.
    
    And even extracting and inserting the code was not stable. Excel will arbitrarily change the casing of “.path” to “.Path” for no reason and add and remove whitespace between functions as it see fit. It was such a pain. We also had a hard time handling unicode strings for instance containing a degree sign. And the list goes on.
    
    source
    PlexSheep@infosec.pub ⁨1⁩ ⁨year⁩ ago
    Perhaps M$ does that specifically to make it hard to work with their formats? That way, tools like libre office stay not 100% compatible, preserving their market share.
    
    source
    -> View More Comments
- oce@jlai.lu ⁨1⁩ ⁨year⁩ ago
  Just fork git to handle zipping, formatting and ignoring metadata! Or just put your office document in the cloud and use the basic versioning it provides.
  
  source
MacNCheezus@lemmy.today ⁨1⁩ ⁨year⁩ ago
Doesn’t matter, to git they are still binary files, which means it’ll check in each revision as an entirely new copy.

source
- Zagorath@aussie.zone ⁨1⁩ ⁨year⁩ ago
  Someone could probably build a tool which sits in between you and Git, which unzips the file before committing and after pulling, so Git sees the raw xml file, but you always see the zipped docx.
  
  source
  - petersr@lemmy.world ⁨1⁩ ⁨year⁩ ago
    Yeah, I made such a tool - and kept polishing edge cases until I gave up.
    
    source
  - MacNCheezus@lemmy.today ⁨1⁩ ⁨year⁩ ago
    I’m sure you could, but yes, it’s likely not worth the trouble.
    
    source
  - refalo@programming.dev ⁨1⁩ ⁨year⁩ ago
    a pre hook filter that beautifies and sanitizes the xml should fix that
    
    source
  - uis@lemm.ee ⁨1⁩ ⁨year⁩ ago
    git-scm.com/…/Customizing-Git-Git-Attributes#filt…
    
    source
- JackbyDev@programming.dev ⁨1⁩ ⁨year⁩ ago
  Which isn’t any different than keeping them as separate files space wise so what’s the problem?
  
  (Other than Word having built-in versioning.)
  
  source
  - MacNCheezus@lemmy.today ⁨1⁩ ⁨year⁩ ago
    
    what’s the problem?
    
    It’s basically just keeping a bunch of separate files but with extra steps.
    
    source
    JackbyDev@programming.dev ⁨1⁩ ⁨year⁩ ago
    I would genuinely rather use git in such a scenario than not because there are plenty of other useful features over a bunch of files in a folder. Sure, obviously if the file is massive it is inconvenient, but that’s not a fair comparison because we’re comparing multiple copies “FINAL FINAL FOR REAL” in a folder anyways. There isn’t suddenly less size that way. It seems incredibly silly to describe it as “keeping files with extra steps” because people aren’t using git for space saving, they’re using it for version tracking. Everything git does is “keeping files with extra steps.”
    
    source
    -> View More Comments
- uis@lemm.ee ⁨1⁩ ⁨year⁩ ago
  I think you can write clean/smudge filter that will turn docx into tree(folder)
  
  source
  - MacNCheezus@lemmy.today ⁨1⁩ ⁨year⁩ ago
    You probably can but here’s why that’s still not gonna be all that effective.
    
    source
    uis@lemm.ee ⁨1⁩ ⁨year⁩ ago
    Still better than having 30 copies of same document and forgetting which was the last one.
    
    source
    -> View More Comments