In the case of Assange, Australian politicians did a lot of work to get him released in the end[1], including:
* Sending a delegation representing all major political parties to the US to argue for the release of Assange. Imagine picking the Republican and Democrat politician LEAST likely to want to cooperate on anything, and those two would have been Australia's equivalent representatives in this delegation. Reports afterwards of the meeting at DOJ HQ indicated it wasn't the type of meeting where the Australians would have brought Tim Tams to share around the room.[2]
* The Australian parliament voted publicly 2:1 on a motion for Assange's release.
* Repeated petitioning through ambassadors in the UK and US, official visits of Australian politicians, etc. Not in private either, as is typically the case for diplomatic affairs.
* Australian politicians attending UK extradition hearings.
* After getting agreement to a plea deal, flying Australian ambassadors for the UK and US to the court of a one-pub-town in the middle of the Pacific Ocean no one has heard of (Northern Mariana Islands) in support of Assange, then all of them flying back to the Australian prime minister's aircraft terminal for a welcome home bevvy.
This was all at a time too where "Free Assange" posters and graffiti was _widely_ distributed across Australian cities.
> authors won’t publish if anyone can republish their work for free
There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).
[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.
The people who publish are people who have reason to publish when they can be copied. Typically either they have already been paid, or they expect to gain market share by being free.
People who need renumeration to continue working will not publish.
Intellectual property rights, as much as I dislike the RIAA and MPAA, created a way for more players to enter the market, because it created a way for their needs to be met.
"Sweat of the brow" doctrine has been rejected in most countries.[1] Even Europe's Database Directive, probably the closest thing to an implementation of this doctrine, largely doesn't do much in practice.
An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2]
AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).
If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:
1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?
2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?
3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.
"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.
I wasn't aware of the "sweat of the brow" doctrine until you mentioned it. But you seem to be using it incorrectly. The doctrine only states that creativity or originality isn't required to make a work copyright-able.
Even if this were an accepted principle, that wouldn't change the principle of free use. In all of your examples, only re-printing all or substantial portions of the books of Andre D. Short would be copyright violations. Just referencing facts from Short's books, or even including small quotes, in your own new work is not a violation.
The parent comment I replied to is concerned with "life's work got appropriated without consideration, compensation or consent". To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions. Today in most jurisdictions copyright laws do not care the slightest about an LLM ingesting databases -- phone directories, sport fixtures and results, someone's life work measuring the dimensions of frogs, etc. 100% of the original factual data could be learned by the LLM, and 100% could be output all at once.
^ Of course there are other ways to alleviate the concerns too such as universal basic income, government grants, etc for someone who wants to dedicate their life to measuring the dimensions of frogs, or whatever else their interest may be. There would however be some geopolitical/trade issues involved--a population would have to be comfortable doing the heavy lifting only to have another country simply use the work freely and instead dedicate their lives to something less favourable such as building missiles.
> To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions
No, it wouldn't. "Sweat of the brow" applies to collections of facts whose compilation required effort. "Life's work" is a superset of that. Originality and creativity, which are required to copyright something, are also work.
The US government's official position on LLMs is (very simply paraphrased) that LLMs are sufficiently transformative and do not hamper the potential market of authors of training material, therefore, copyright claims arising from training material should not be successful.[1] For original and creative training material, for example, a Harry Potter novel, seemingly the US government is asking the courts to set aside some previous questionable findings such as copyright existing very loosely in the likeness of fictional characters (impacting the likes of fan fiction). Can a human -- or LLM -- create a story about children travelling on a train from New York to a school of magic in the "wild west", with many loose similarities to Harry Potter for those familiar with those books? The US government appears to be saying this is OK, especially with the view that the market for Harry Potter is not diminished by a "wild west magic school" book in its likeness.
However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.
> LLMs are sufficiently transformative and do not hamper the potential market of authors of training material
And OP's contention is obtaining the training material and using it in training requires making unauthorized copies. That's the infringement; training, not inference.
Furthermore inference indirectly affects the market for the artist's future work. Don't need the writers and artists the LLM trained on anymore, when it can do similar work for free.
How long would it then take to be able to use the backed up data? Wait for a war to end and a replacement data centre to be built...? By that time most data probably no longer matters (e.g. business no longer exists).
It's more likely the entire data centre (not just backups) would need to be built underground (or cut and cover) at enormous expense. A price that perhaps for certain data sovereignty reasons the government of Bahrain (or companies in Bahrain requiring it) would be happy to pay?
Another way to do things on the cheap could be small-scale "covert hosting". Buy an apartment or house, maintain it to give an outside appearance of being an apartment or house, but inside it has a few racks of IT equipment. This has been done in the past for telephone exchanges in some countries, not for security reasons, but rather to hide an ugly bit of infrastructure that due to technology limitations of the time had to be located deep within a residential neighbourhood.
Global offline maps including public transport routing and timetables is a fantastic (and if I'm not mistaken, also unique) feature. Google Maps' offline feature by contrast doesn't work globally (some areas are excluded), doesn't include public transport, etc.
Keeping in that same theme--I've love to see Mozilla Translations models (or similar) integrated in GNOME apps where it makes sense to do so--such as Document Viewer, for good-enough offline text translation built into every GNOME environment.
Such features are highly user visible and set GNOME apart from competition in ways that are easy to describe to anyone. No EULAs, no data sovereignty concerns, no dark patterns--just features ready to go for users out-of-the-box that are incredibly useful, user-friendly and consistent.
In saying that--there's of course all the great GNOME stuff under the hood for geeks too--support for almost every audio and video format to have ever existed easily enabled, glycin, sandboxing of applications, easy and consistent software package management for everything (now with obsolescence management built in too)...etc.
Translation within these apps is generally not an easy feature to implement. Do you automatically try and detect the language and suggest a translation? Do you make language translation controls obvious and put them front and centre, or hide them away assuming they're infrequently used? Do you show translations side-by-side, line-under-line, or just replace the original text with translated text? Then on more complex matters such as papers, do you try and preserve original formatting (very hard, particularly for things like tables), do you accommodate words split across lines, etc.
There does exist https://flathub.org/en/apps/dev.ters.LocalTranslate as a standalone offline text translation application for GNOME desktop environments but a user would have to copy+paste text between applications to use it.
I wasn't thinking of the UI being "highlight text -> context menu -> translate", rather, one of:
1. Open bonjour.md and a banner automatically appears above the document asking whether you'd like to translate from French -> German (assuming the desktop environment language was set to German). This is more or less how Firefox handles offline translation of web pages.
2. Open bonjour.md and click a translate button, then get asked which language pair to choose from.
I can however see how the "highlight text -> context menu -> translate" UI pattern may make sense in multi-language settings, such as a chat room, or web browser where text in multiple language may be presented on the same page. For a terminal session however, it may be better to require the user to pipe text to a translation command line utility where possible--but there are exceptions to this too such as ncurses interfaces.
For me-south-1 (Bahrain), all 3 data centres providing the redundancy were blown up by Iran.[1] The redundancy was localised to small geographic area and a single government--something customers of AWS were hopefully aware of when they entrusted AWS with their data.
It's always buyer beware for any claims of availability. Engineers completing a FMECA[2] will (or should) always state upfront what type of failure modes they've deliberately excluded (such as meteor strike) or else every FMECA would be full of failure modes that have never been measured, and are not worth anyone's time worrying about. These exclusions vary by application--a time capsule, seed vault, etc are intended to outlast wars and collapses of empires. Typically a bunch of data centres aren't designed to withstand such failures.
I do think however it'd be reasonable to include the prospect of war for calculating data centre / cloud service availability. Especially in a place such as Bahrain where the country is obviously concerned enough about the prospect of war to have built very permanent and expensive air/missile defence sites. New Zealand on the other hand--maybe not so important to consider.
Regions are really the scale of disaster isolation only in extreme cases - such as global catastrophe (meteor strike taking out a city) or in this case, when actively targeted in war. I don't really see the same thing happening to a US or European region.
As far as I know, the attacks happened at different times. If Amazon knew that they had lost some data redundancy, shouldn’t they have been quickly mirroring that out of the region?
"You choose the AWS Region(s) in which your content is stored. You can replicate and back up your content in more than one AWS Region. We will not move or replicate your content outside of your chosen AWS Region(s) without your agreement."
This. We have (well, had) customers running in me-south-1 and once the first AZ went down we wanted to proactively move their data to other regions even just as cold backups. But our legal department slapped that down pretty quickly.
Most likely, their own data residency terms prohibit this. It would be interesting to know if, when 2 out of 3 AZs got destroyed, customers got a heads up to move their data to a different region?
We received repeated, constant heads up to move our data by the first AZ much less second. The problem is that nobody is storing data in Bahrain unless there are data residency requirements for it.
nobody wakes up one morning and chooses to launch instances, CDN or S3 and would choose Bahrain as that without a requirement to, we were contractually and legally forbidden (in the middle as a vendor) to copy even encrypted data where we don't have the key out for redundancy, so the best we could do was tell our subcustomers to download all of their buckets to their office or some employee laptops at their office
robots.txt was only intended to help search index crawlers not get stuck in endless crawl loops for badly designed websites.
What you suggest is explicitly not a purpose of robots.txt per RFC9309[1]:
"These rules are not a form of access authorization."
HTTP 429 and HTTP 403 are what servers are meant to return to clients to slow them down or tell them to stop doing something without having first gained authorisation.
Wouldn't it be possible to confirm who may be supplying imagery by taking some IRGC supplied imagery such as [1] and figure out which satellite(s) were overhead at the time and capable of taking the image? It may not be particularly straightforward to do so--but seems theoretically possible with:
- Consideration of time-of-day/azimuth of overhead satellite based on shadows and perspective.
- Comparison to other imagery of the same location checking for portable objects (such as vehicles) and keeping track of when they appear and disappear, and whether they're present in the imagery being analysed.
- Consideration of cloud coverage and weather conditions to rule out unwanted satellite passes.
Then write up a report and publish it, making it a believable and concrete finding instead of a finding that otherwise may just be perceived to be speculation/vibes.
A description of how GFP was first extracted is at [1] and includes a custom-made "jellyfish cutting machine" capable of processing 600 rings an hour, and harvesting and processing 50,000 jellyfish to create just a single milligram of purified aequorin, with multiple milligrams needed to conduct experiments required to isolate/describe the compounds of interest. It took the scientists 5 years to reach the point where they had discovered the chemical structures and understood what they'd extracted from the goop of 50,000+ jellyfish.
A video tutorial at [2] then explains what to do next once GFP has been obtained from commercial sale and use of modern cloning, to skip the need to own a jellyfish cutting machine. E.g. need an antibiotic to extract the 1% of e.coli which will end up being modified vs. 99% of e.coli that won't be changed by presence of GFP.
To the point of the DavidRBellamy X thread, isn't the suggestion that an LLM wouldn't know whether it's worthwhile extracting some substance from the stomach of a crocodile, or from the beak of a penguin that only lives in Antarctica? And that much of the difficulty rests with hard physical work of the equivalent of cutting 600 rings an hour off jellyfish, for 50,000+ jellyfish, and conducting experiments to isolate and figure out the use of an extracted compound. Or ordering some obscure chemical compound from a supplier that is normally produced at a rate of 1mg/year and instead the LLM wants to order 1kg all of a sudden, raising eyebrows about why someone might be making the order.
The Twixxer thread seems to be arguing that an AI can't do research into novel pathogens without people, which is probably true. I haven't met anyone who is worried about that and I think it's either mistaken or a strawman.
I've mostly heard "what happens if the next ISIS or the next Jim Jones can modify ebola to be waterborne or create super-MRSA because AI simplifies the process?" which is why I ask what it would take to create a more virulent virus on a shoestring budget. Not do novel Nobel-worthy research, just copy some existing paper from ten years ago.
I couldn't find a model/calculator that would help visualise typical COP values for particular climatic conditions. However, you'll find in your travels that a COP of ~2-2.5 is typically achieved for ambient (outdoor) temperatures of -15oC (for air sourced heat pumps).
If 300L of water at 15oC is filled into a tank and needs to be heated to 60oC within 2 hours during ambient temperature of -15oC, you get very approximately (no thermal losses considered):
- An output energy need of approximately (4190300(60-15))/(60*120)=~8kW (56MJ/2h)
- An input electricity need of approximately (8/2.2)=~3.6kW
Instead of a resistive heating hot water unit requiring 16kWh to do this job, you could use a heat pump hot water unit requiring 7.2kWh, cutting electricity use in half.
And this is for arguably the most extreme use case for a heat pump hot water unit where it's "cold started" right at the coldest moment in Winter in cool-temperate climates (such as SE Australia). Think for example, arriving at a ski chalet and having to turn on the hot water unit before someone can take the first hot shower.
On a more typical day of the year, perhaps with overnight ambient temperature of 10-15oC, the COP would rise to ~4, equating to an electricity consumption of 2kWh to heat the 300L of water. A lot of units will be set to heat during the warmest part of the day, let's assume an ambient temperature of 25-30oC, where a COP of ~5-6 is more typically achieved. However, there are obviously diminishing returns for COP of 4 vs 5.
In arctic climates, heat pumps are still used, but with a ground or aquifer source rather than ambient air source.[2]
* Sending a delegation representing all major political parties to the US to argue for the release of Assange. Imagine picking the Republican and Democrat politician LEAST likely to want to cooperate on anything, and those two would have been Australia's equivalent representatives in this delegation. Reports afterwards of the meeting at DOJ HQ indicated it wasn't the type of meeting where the Australians would have brought Tim Tams to share around the room.[2]
* The Australian parliament voted publicly 2:1 on a motion for Assange's release.
* Repeated petitioning through ambassadors in the UK and US, official visits of Australian politicians, etc. Not in private either, as is typically the case for diplomatic affairs.
* Australian politicians attending UK extradition hearings.
* After getting agreement to a plea deal, flying Australian ambassadors for the UK and US to the court of a one-pub-town in the middle of the Pacific Ocean no one has heard of (Northern Mariana Islands) in support of Assange, then all of them flying back to the Australian prime minister's aircraft terminal for a welcome home bevvy.
This was all at a time too where "Free Assange" posters and graffiti was _widely_ distributed across Australian cities.
[1] https://en.wikipedia.org/wiki/Julian_Assange#Plea_bargain_an...
[2] https://www.abc.net.au/news/2024-06-27/inside-the-us-austral...
reply